Kamal Hammouda, Rishi Dakarapu, Chathruckan Rajendra, Hyojeong Lee, Sahar Almahfouz Nasser, Sharmistha Rudra, Vasantha Kolachala, Caitlin Diefendorf, Irina Geiculescu, Sushma Maddipatla, Anant Madabhushi, Subra Kugathasan
This novel AI-based framework predicts MES across geographically diverse adult and pediatric UC datasets. Its strong performance, including comparability to expert gastroenterologists in EPC, supports its potential as a decision-support tool for standardized endoscopic monitoring. However, prospective multicenter validation is warranted prior to routine clinical implementation.
BACKGROUND: Endoscopic assessment in ulcerative colitis (UC) is critical for clinical decision-making but limited by interobserver variability. AI-powered systems may improve consistency, but dataset heterogeneity has hindered clinical translation. We developed and evaluated a deep learning framework for standardized Mayo Endoscopic Score (MES) prediction across adult and pediatric populations from diverse regions.
PATIENTS AND METHODS: Using three multi-institutional datasets, we trained and validated convolutional neural network models to predict MES. The primary dataset (LIMUC; 564 adults, 11 276 images, Turkey) supported model development, with external validation on TMC (308 adults, 7978 images, China) and the Emory Pediatric dataset (EPC; 80 children, 113 images, USA). Two models were developed: UC-Re for binary remission classification (MES 0-1 vs 2-3) and UC-MES for four-class grading (MES 0-3). Imaging artifacts were corrected using inpainting, and Fourier-Spatial Image Harmonization (FSIH) mitigated inter-institutional domain shifts. Performance was evaluated using area under the operating characteristic curve (AUC), F1-score, and quadratic weighted kappa (QWK).
RESULTS: UC-Re achieved AUCs of 0.98, 0.95, and 0.98 across LIMUC, TMC, and EPC, with F1-scores of 0.92, 0.87, and 0.93, respectively. UC-MES demonstrated strong ordinal consistency (QWK = 0.81-0.85), comparable to inter-expert agreement (QWK = 0.88). Most misclassifications occurred between adjacent MES categories, reflecting human-like patterns.
CONCLUSIONS: This novel AI-based framework predicts MES across geographically diverse adult and pediatric UC datasets. Its strong performance, including comparability to expert gastroenterologists in EPC, supports its potential as a decision-support tool for standardized endoscopic monitoring. However, prospective multicenter validation is warranted prior to routine clinical implementation.