Xiang Yong, Zhen Li, Qiang Wu, Huaiyuan Hu
Available cohort-level evidence suggests high sensitivity and strong overall discrimination for colorectal cancer detection on histopathological whole-slide images. However, the exploratory estimates are dominated by one development program and should not be interpreted as a mature, independently replicated, multisource evidence base. Specificity remains vulnerable to endpoint heterogeneity and slide-level calling rules. Current evidence supports the use of deep learning as an adjunct for prescreening, triage, or quality control, rather than as a standalone diagnostic replacement.
BACKGROUND: Deep learning systems are increasingly used for colorectal cancer detection in digital histopathology, but studies often report heterogeneous endpoints, patch-level metrics, or non-thresholded area under the curve (AUC) values. We estimated the diagnostic accuracy of clinically interpretable slide-level or patient-level models.
METHODS: We searched PubMed, Embase, Scopus, Web of Science Core Collection, and IEEE Xplore for deep learning studies of human colorectal histopathology or whole-slide images. Eligible studies provided extractable or reconstructable slide-level or patient-level true-positive (TP), false-positive (FP), false-negative (FN), and true-negative (TN) values for colorectal cancer, colorectal adenocarcinoma, or closely aligned malignant colorectal histopathology detection. Risk of bias was assessed using QUADAS-2 with AI-pathology-specific considerations. The strict primary synthesis excluded endpoint-caveat cohorts. Summary sensitivity and specificity were estimated using a bivariate random-effects Reitsma model. A second reviewer verified eligibility, 2 x 2 tables, and QUADAS-2 judgments.
RESULTS: The search identified 4,687 records; 2,307 unique records were screened. Three source studies contributing 14 validation cohorts met the inclusion criteria for the strict primary synthesis; five additional endpoint-caveat studies contributing seven cohorts were retained only for expanded and sensitivity analyses. Primary cohorts included 5,402 TP, 439 FP, 68 FN, and 3,596 TN. The exploratory bivariate model, which treated cohorts as independent observations, estimated a sensitivity of 0.982 (95% CI, 0.976-0.986), a specificity of 0.959 (95% CI, 0.922-0.979), and an AUC of 0.979. Wang (2021) contributed 12 of 14 cohorts; within-study correlation could therefore make these confidence intervals overly precise. Sensitivity was stable across scenarios, whereas crude specificity varied with endpoint definition, analysis unit, and large true-negative denominators.
CONCLUSION: Available cohort-level evidence suggests high sensitivity and strong overall discrimination for colorectal cancer detection on histopathological whole-slide images. However, the exploratory estimates are dominated by one development program and should not be interpreted as a mature, independently replicated, multisource evidence base. Specificity remains vulnerable to endpoint heterogeneity and slide-level calling rules. Current evidence supports the use of deep learning as an adjunct for prescreening, triage, or quality control, rather than as a standalone diagnostic replacement.