Faiza Tariq, Hira Hameed, Tariq Malik, Michael S. Mollel, Muhammad Ali Imran, Qammer H. Abbasi
Glaucoma is a silent and progressive eye disease that can cause irreversible vision loss if left undiagnosed. Early detection and screening are therefore critical, particularly in underserved regions with limited ophthalmologic resources. This paper presents a deep learning based hybrid ensemble framework for early glaucoma detection from retinal fundus images. The proposed model integrates ResNet50 and Vision Transformer (ViT) architectures, enabling one to capture local spatial features and the other to learn global contextual representations. Preprocessing involves extracting the green channel and applying contrast-limited adaptive histogram equalization (CLAHE) to enhance retinal vessel visibility and texture detail. To improve generalization and mitigate class imbalance, data augmentation is applied before model training. The ensemble fuses softmax outputs from both models to produce the final classification. Experiments conducted on the standardized SMDG-19 dataset comprising 17,242 labeled fundus images demonstrate that the ensemble achieves 99.8% accuracy, F1-score = 1.0, and MCC = 0.994, confirming strong robustness to contrast variations, noise, and illumination changes.