Krittee Chesdachai, Kullathorn Thephamongkhol, Anucha Chaichana, Jiraporn Setakornnukul
A properly specified clinical model was a strong benchmark for radiomics. Single-best pipeline results can overstate radiomic value in small cohorts, and biascontrolled evaluation with external validation is essential before clinical use.
BACKGROUND AND PURPOSE: Non-complete response after chemoradiotherapy in unresectable locally advanced head and neck squamous cell carcinoma (LA-HNSCC) has been associated with poor prognosis. We compared clinical, computed tomography (CT), and magnetic resonance imaging (MRI) radiomics models for predict non-complete response of the primary tumor, and estimated each data type's marginal contribution across all configurations.
MATERIALS AND METHODS: This retrospective study included 88 patients (67 complete and 21 non-complete responders) treated with definitive chemoradiotherapy. Radiomic features were extracted from pre-treatment contrast-enhanced CT and multi-parametric MRI. Models were optimized under nested leave-one-out cross-validation across 366 configurations spanning imbalance handling, feature selection, and classifier. Each data type's marginal contribution was estimated as the mean paired AUC difference versus a parsimonious clinical model. Early and late fusion were tested, and clinical utility was assessed by exploratory decision curve analysis.
RESULTS: In a best-configuration comparison, the MRI model achieved the highest discrimination (AUC 0.91, 95% CI 0.81-0.98) but did not significantly outperform the parsimonious clinical model (AUC 0.75, DeLong p = 0.058). Across all configurations, no single-modality or feature-level fusion exceeded the clinical baseline, each performing significantly below it (all p < 0.05). Only decision-level fusion combining clinical and imaging data was positive, reaching significance for clinical plus CT (mean ΔAUC +0.059, p = 0.031).
CONCLUSIONS: A properly specified clinical model was a strong benchmark for radiomics. Single-best pipeline results can overstate radiomic value in small cohorts, and biascontrolled evaluation with external validation is essential before clinical use.