A. Hiras, A. Jayaraman, A. Gadari, A. S. Mankumare, A. Mynampati, J. K. Chhablani, S. C. Bollepalli, K. K. Vupparaboina
Automated segmentation of Optical Coherence Tomography (OCT) images is a critical component of structural biomarker extraction for retinal diagnostics. Deep learning models achieve state-of-the-art performance on controlled datasets, yet exhibit unpredictable failures on real-world data. Current quality gates rely on device-reported scan quality scores, which have been shown to be unreliable predictors of segmentation performance. We define scan quality in a task-specific sense, that is, whether a given B-scan will yield a reliable segmentation from a particular trained model. Under this definition, we systematically evaluate No-Reference Image Quality Assessment (NR-IQA) metrics, general-purpose pretrained representations, and domain-specific pretrained representations as alternative quality gates. We use choroid segmentation as the prototype task, with a dataset of 6,076 OCT B-scans from 80 subjects. These quality gate candidates are evaluated at three levels: scalar metrics (BRISQUE, NIQE, PIQE, SNR, PSNR), supervised linear probing, and unsupervised partitioning (K-Means) of feature vectors and learned representations. All scalar NR-IQA metrics proved inadequate (|r| < 0.20). General-purpose pretrained representations (EfficientNet-B0, ResNet-50, ViT-B/16) outperform NR-IQA, achieving ROC-AUC up to 0.77, indicating that learned representations are better suited to task-specific quality gating than hand-crafted scalar statistics. Retinal foundation models (FMs) further improve performance: RETFound, an OCT-specific FM, achieves ROC-AUC of approximately 0.81. Unsupervised K-Means partitioning of the embeddings recovers the same hierarchy geometrically: only the two retinal foundation models produce quality-aligned clusters that exceed a patient-level permutation null, while NR-IQA and general-purpose pretrained feature spaces do not, suggesting that domain-specific pretraining provides an additional benefit beyond general-purpose learned representations.