Muhammad Uzair Gill, Parvathy Rajendran
Small object detection in high-resolution aerial images is difficult because of the objects limited pixel size, sparse features and the intricacy of the backgrounds. To address this problem, we introduce AWAGS-YOLO (Attention module, Weighted feature fusion bi-directional feature pyramid network, Attention module, Global attention module, Scylla Intersection over Union-You Only Look Once), an improved one-stage detector based on You Only Look Once-11 (YOLO-11) framework which enhances small object detection. To better capture fine object features in cluttered environments, the AWAGS-YOLO framework integrates attention mechanisms into the backbone, employs a bi-directional feature pyramid network with learnable weighted feature fusion in the neck, and incorporates global attention modules and a geometry-aware loss function. Using these advances, our model achieves much greater accuracy for small object identification than the standard YOLO-11. On the VisDrone2019 benchmark, AWAGS-YOLO outperforms the YOLO-11 baseline by 7.46 % in mean Average Precision (mAP50:95). On our custom Realm aerial dataset, it achieves an 8.24 % mAP50:95 advantage over standard YOLO-11 variant. These findings show that our focused enhancements efficiently address the stated problem, resulting in superior detection performance for small objects in aerial images while maintaining real-time efficiency. However, challenges such as improving recall for extremely small or occluded objects remain, indicating directions for future work.