Lydia A Schoenpflug, Nikki J van den Berg, Sonali Andani, Nanda Horeweg, Jurriaan Barkey Wolf, Tjalling Bosse, Viktor Koelzera, Maxime W Lafarge
Sensitivity to staining variation remains a major barrier to deploying computational pathology (CPath) models as hematoxylin and eosin (H&E) staining varies across labs, requiring systematic assessment of how this variability affects model prediction. Yet, current evaluation practices rarely attribute performance changes to specific, measurable staining properties, limiting interpretable robustness assessment under controlled conditions. We introduce a three-step protocol that explicitly simulates H&E staining variation toward known, measurable targets, supporting controlled assessment of model performance shifts. Step 1: Select reference staining conditions, Step 2: Characterize test set staining properties, Step 3: Apply CPath model(s) under simulated reference staining conditions. Here, we construct a new reference staining library based on the PLISM dataset. As a use case, we evaluate 306 microsatellite instability (MSI) classification models on the SurGen colorectal cancer dataset (n = 738), including 300 attention-based multiple instance learning models trained on TCGA-COAD/READ across three feature extractors (UNI2-h, H-Optimus-1, and Virchow2), and six public models. Classification performance was measured as AUC, and robustness as the min-max AUC range across four simulated staining conditions (low/high H&E intensity, low/high H&E color similarity). Across models and staining conditions, classification performance ranged from AUC 0.769-0.911 (Δ = 0.142). Robustness ranged from 0.007 to 0.079 (Δ = 0.072), and showed a weak inverse correlation with classification performance (Pearson r = -0.28, 95% CI [-0.39, -0.17]). Thus, using MSI classification in colorectal cancer as a use case, we demonstrate that the proposed evaluation protocol enables robustness-informed CPath model selection and provides insight into how H&E staining conditions affect model performance, with the potential to support reliable deployment across clinical settings. Code and reference library are publicly available at https://huggingface.co/datasets/CTPLab-DBE-UniBas/staining-robustness-evaluation and https://github.com/CTPLab/staining-robustness-evaluation.