Jiayao Chen, Xiaozhou Feng, Hao Hu, Mengyan Liu, Qiuhe Ji, Hui Guo, Wenhua Hu, Fei Xie
This review examines retinal image analysis from the perspective of clinical translation rather than benchmark-oriented model comparison. Instead of organizing prior work solely by disease category or model family, we synthesize recent advances through four connected dimensions: imaging modality, task taxonomy, methodological paradigm, and translational bottleneck. We compare major tasks, including classification, detection, segmentation, grading, progression prediction, and treatment-response assessment, across fundus photography, OCT/OCTA, and angiographic imaging. We further review the roles of CNNs, U-Net variants, 3D models, Transformers, graph-based methods, hybrid architectures, foundation models, self-supervised pretraining, and multimodal learning under different data conditions and clinical constraints. Beyond technical progress, we analyze why many high-performing systems still fail to translate reliably into practice, highlighting challenges related to distribution shift, label inconsistency, limited external validation, image-quality control, calibration, fairness, and workflow integration. We argue that the next stage of retinal AI requires transferable representations, robust external evaluation, trustworthy uncertainty handling, and demonstrated value within real-world care pathways, rather than incremental architectural novelty alone.