Daniel Curto, Susana Hernandez, Marta Alonso, Melina Peressini, Ricardo Garcia-Lujan, Pablo Gamez, Jose Luis Campo-Cañaveral de la Cruz, Wei Sun, Shuang Han, Dongmei Lin, Christiane Kümpers, Christian Matek, Felix Elsner, Konrad Steinestel, Javier Martin-Lopez, Clara Salas, Virginia Calvo, Mariano Provencio, Luis Paz-Ares, Jon Zugazagoitia, Fernando Lopez-Rios, Esther Conde
AI-based %RVT scoring was concordant with expert pathologist consensus and consistent across real-world, multicenter, multi-scanner cohorts, with most algorithmic errors mirroring those encountered in human practice. By providing explainable %RVT grounded in international recommendations, our framework may facilitate harmonization of PR assessment for perioperative therapies.
BACKGROUND: Pathologic response (PR), expressed as the percentage of residual viable tumor (%RVT), has emerged as a surrogate endpoint after neoadjuvant chemoimmunotherapy in non-small cell lung cancer (NSCLC). However, PR assessment remains subject to interobserver variability and is insufficiently standardized across institutions and clinical trials. We hypothesized that artificial intelligence (AI)-based %RVT scoring would be concordant with expert manual assessment and consistent across real-world, multi-scanner cohorts.
MATERIALS AND METHODS: Four retrospective NSCLC cohorts treated with neoadjuvant chemoimmunotherapy and surgery were analyzed. The internal cohort included 38 patients; three external cohorts from China, Germany, and Spain contributed 97 additional patients. Weighted consensus manual %RVT served as the reference standard for benchmarking an AI algorithm. Discordant cases were interrogated for algorithmic bias.
RESULTS: Overall, 135 patients and 1,344 primary-tumor H&E slides were evaluable. Concordance between AI-based and manual assessment was at least moderate across all cohorts. For major pathologic response (MPR) classification, the AI algorithm achieved a pooled accuracy of 91% (95% CI, 85-95), sensitivity of 90% (82-95), specificity of 93% (81-99), and an area under the curve ≥0.95 in all cohorts. Discordant classifications occurred in 8.9% of cases and were mainly attributable to stromal underrepresentation and misclassification of non-neoplastic elements.
CONCLUSIONS: AI-based %RVT scoring was concordant with expert pathologist consensus and consistent across real-world, multicenter, multi-scanner cohorts, with most algorithmic errors mirroring those encountered in human practice. By providing explainable %RVT grounded in international recommendations, our framework may facilitate harmonization of PR assessment for perioperative therapies.