Y. Riyazifar
Lumbar foraminal stenosis is a spatially localized and ordinal MRI interpretation problem: a useful computational system must first identify informative sagittal slices and foraminal regions before assigning severity. We present a retrospective, leakage-controlled multiscale deep-learning study using the public LSS MRI AISSLab cohort of 500 multi-scanner sagittal T2-weighted lumbar MRI examinations. A single frozen 70/15/15 patient partition was propagated across mid-sagittal anatomy segmentation, 2-D/2.5-D slice selection, anchor-free foraminal localization, four-grade region-of-interest (ROI) classification, radiomics, uncertainty analysis, and an exploratory whole-volume 3-D classifier. The selected U-Net achieved mean Dice 0.950 across five foreground anatomical labels (0.956 across all six released labels). The 2.5-D slice selector achieved held-out ROC AUC 0.926 (95% patient-clustered CI, 0.909-0.941). A threshold locked only on the tuning set yielded test sensitivity 0.876 (0.832-0.919) and specificity 0.841 (0.813-0.869). Among 68 test patients with at least one annotated slice, an annotated slice appeared within the top three ranked slices in 68/68 patients (100%; exact 95% CI, 94.7-100%). The detector achieved localization AP50 0.530 but AP75 0.046, while 29.1% (26.0-32.0%) of annotation-negative test slices generated at least one prediction, identifying precise localization as the principal bottleneck. On expert-defined ROIs, a class-weighted scratch CNN achieved quadratic weighted kappa (QWK) 0.638 (0.550-0.706), with 92.7% (90.4-94.8%) of predictions within one grade. Moderate-or-worse and severe AUCs were 0.893 and 0.912, respectively. Compared with a 29-feature radiomics-SVM baseline, the CNN improved balanced accuracy by 0.205, macro-F1 by 0.173, and QWK by 0.360 using paired patient bootstrap. In a secondary uncertainty analysis, mean segmentation entropy strongly tracked mean surface error (Spearman rho = 0.807, 95% CI 0.689-0.881). In contrast, the whole-volume 3-D CNN achieved AUC 0.639 (0.496-0.765) with poor calibration. The findings support an anatomically constrained, localization-aware strategy and demonstrate why raw accuracy or whole-volume classification alone can be misleading in highly imbalanced foraminal stenosis assessment.