科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ PloS one2026-01-01

Compact vision models match domain-specific foundation models for several retinal imaging classification tasks: A systematic benchmark.

Dávid Isztl, Tahm Spitznagel, Gábor Márk Somfai, Rui Santos

原始摘要(英文原文)· Original abstract
Large domain-specific foundation models have been widely adopted for retinal image analysis, yet systematic evidence for their advantage over compact general-purpose architectures remains scarce. We benchmarked nine model configurations spanning 22.8M to 303M parameters (vision transformers, hierarchical Swin Transformers, ConvNeXt, and the domain-specific RETFound models) across four tasks: OCT 8-class disease classification, and three fundus photography tasks (DME severity, glaucoma detection, and DR severity grading). All models were evaluated under identical training conditions, with both pretrained (on natural-domain image datasets) and from-scratch initializations compared using Mann-Whitney U tests. Pretraining improved accuracy by 5.18-18.41 percentage points across all tasks (p < 0.05 throughout), with larger benefits for CFP modalities and harder tasks. Compact hierarchical models (27-29M parameters) matched or exceeded larger architectures on three of four tasks. For instance, the SwinV2-tiny architecture ranked first on OCT, DME, and GL classification. The domain-specific RETFound model (303M) achieved the highest accuracy only on the most challenging task (DR severity grading, where the most severe class is underrepresented at 8% of images), where it outperformed the best compact model by 1.54 percentage points. These results indicate that compact general-purpose models may be sufficient for most retinal classification benchmarks, and that domain-specific foundation models may add higher value mainly for severity grading tasks with skewed class distributions.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related