Eleftherios P. Mamounas, Ming Chen, Joseph A. Sparano, Md Ashequr Rahman, Yating Cheng, Victoria Wang, Robert J. Gray, Priya Rastogi, Eghbal Amidi, Charles E. Geyer, Tommy Boucher, Tanner J. Freeman, Mohammadreza Ramzanpour, Mukund Varma, Hassan Ghani, Caleb Cheng, Casey Bales, Jennifer R. Ribeiro, Hanna Bandos, Nicolas Stransky, Mark R. Miglarese, Matthew J. Oberley, David Spetzler, Milan Radovich, George W. Sledge, Norman Wolmark
A, Overview of the M3T model workflow. The model was developed using the NSABP B-42 translational cohort (N = 2,271) with five-fold cross-validation, integrating H&E image–derived features and selected tabular clinicopathologic features through the M3T (multimodal, multitask transformer). An auxiliary task predicting the lowest BMD T-score enhanced prognostic performance and improved stratification of patients with differential ELT benefit but was not required for model inference. The model outputs a continuous risk score that is dichotomized into low- and high-risk groups. Independent external validation was performed in the TAILORx translational cohort (N = 4,300). Prognostic, predictive, and calibration analyses were conducted to evaluate model performance. B, End-to-end image preprocessing, feature extraction, model training, and output architecture. Tissue regions were detected from H&E-stained WSIs, tiled at 10× magnification, and processed using CTransPath to generate tile-level image embeddings. Tile-level image embeddings and clinicopathologic features were projected into a shared representation space and processed through transformer encoder and decoder modules. During training, an auxiliary task predicting the lowest BMD T-score was incorporated to improve prognostic performance and stratification of patients with differential ELT benefit. The primary task generated a continuous risk score for late DR risk stratification. The dashed box shows the detailed architecture of the transformer encoder (left) and transformer decoder (right).