Zhiming Xiong, Meiting Yu, Tong Lv, Ruixuan Xu, Chongjun Huang, Guohui Yuan, Zhuoran Wang
Fundus diseases are a leading cause of irreversible blindness worldwide. Existing Color Fundus Photography (CFP) and Optical Coherence Tomography (OCT) fusion methods face two major challenges: feature noise interference and the semantic gap between heterogeneous modalities, leading to suboptimal fusion and limited diagnostic accuracy. To address these issues, this paper proposes a Multi-modal Hybrid Fusion Network (MHF-Net), which is designed to suppress noise, enhance intra-modal feature representation, and explicitly establish cross-modal semantic correlations. The proposed framework consists of three core components: a heterogeneous dual-stream encoder to extract multi-level pathological features from both CFP and OCT; a Multi-scale Feature Enhancement (MFE) module to reduce background noise while reinforcing intra-modal features; and a Hierarchical Contextual Fusion (HCF) module to explicitly establish deep semantic correlations across modalities and scales. Evaluated on the Topcon-MM dataset, MHF-Net achieves a Macro F1 of 80.91%, a Macro Recall of 76.80%, and a Sample Accuracy of 67.32%, demonstrating competitive performance against state-of-the-art multimodal models and recent vision foundation model baselines. This study demonstrates that explicit hierarchical fusion with multi-scale feature enhancement effectively resolves the semantic gap between CFP and OCT modalities, providing a robust and clinically promising solution for automated multimodal diagnosis of complex fundus diseases.