科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Scientific Reports2026-06-15· Computer science

Explainable AI for diabetic retinopathy detection using vision transformers

Mustafizur Rahaman, Masrufa Akter Muni, Saima Tasnim, Sanjida Shahid Juthi, Rakibul Islam, Khandakar Rabbi Ahmed

原始摘要(英文原文)· Original abstract
The heterogeneous acquisition, variability of orientation, and subtle lesions continue to challenge the screening of diabetic retinopathy through color fundus photographs. We formulate DR grading as a binary triage task (No-DR vs DR) and propose the All-ViT Hybrid framework, integrating complementary pretrained transformer backbones within a stability-oriented training schedule (head-only warm-up, partial unfreezing), optimized using AdamW with OneCycle scheduling, class-weighted cross-entropy, and mixed precision. Preprocessing includes luminance-space CLAHE and retinal field-of-view masking, and data splitting is performed using stratified sampling to preserve class balance. We perform post-hoc threshold tuning and test-time augmentation (optional), which is available for operating-point control and robustness. Over the competitive baselines (ConvNeXt, MobileNetV3-Large, EfficientNet-B0, DenseNet201), All-ViT Hybrid has an Accuracy of 0.9754, F1 of 0.9761, Precision of 0.9658, Recall of 0.9866, and measures of agreement κ of 0.9509, MCC of 0.9511, and Jaccard of 0.9532. Compared to the best baseline, F1 achieves a gain of +1.11 percentage points, an improvement in accuracy of +1.13, and a recall improvement of +0.81 percentage points without compromising controllable precision and these experiments were conducted on the APTOS 2019 Blindness Detection dataset. These visualizations provide qualitative interpretability and are not quantitatively validated against lesion-level annotations due to dataset limitations. These findings suggest that integrating complementary transformer representations within a unified fusion pipeline provides strong and threshold-adjustable performance for DR triage under real-world variability. The framework remains modular and extensible, with potential applicability to larger input resolutions, multi-class grading, and multi-modal clinical settings. The future research program will test external generalization and probability calibration across devices and centers.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Explainable AI for diabetic retinopathy detection using vision transformers — 科研速览 Science Skim