Shahzad Hussain, Iqra Mumtaz, Chong Wang, Pei Lv
Small object detection in aerial imagery is a challenging task due to the minimal pixel information in dense clutter, scale variation, and complex backgrounds. YOLOv9 has demonstrated the effectiveness of Programmable Gradient Information (PGI) in mitigating feature degradation. However, its fully convolutional architecture lacks the capability for global context modeling, which is critical for resolving ambiguities in small targets. To address these limitations, we propose SF-YOLOv9, a hybrid architecture that enhances YOLOv9c by improving the backbone through the integration of a novel PGI-Aware Swin Fusion Block (Transformer-GELAN) at its final stage. This module effectively preserves high-resolution local features while injecting long-range global context through Swin Transformer-based fusion. It results in richer and more discriminative semantic representations. We introduce a Dual-Path Spatial and Channel Attention Module (DSCAM) into the main detection head and the reversible auxiliary branches of PGI. By refining attention across all supervisory signals, DSCAM significantly improves gradient flow and feature fidelity during PGI training, reducing missed detections and false positives. We evaluate SF-YOLOv9 on VisDrone and NWPU-VHR-10 datasets to demonstrate the effectiveness of SF-YOLOv9. It outperformed the baseline models, achieving 49.1% mAP@0.50 on VisDrone and 98.3% mAP@0.50 on NWPU VHR-10 in small-object detection.