Zening Wang, Wenjuan Li, Yangquan Fan, Liyuan Wang, Guoquan Yao
To address the challenges of ship detection in remote sensing images, where ship targets are relatively small and the scene complexity often leads to false detections and missed detections, this paper proposes an improved YOLO11s method called YOLO11s-Adaptive Pyramid Fusion Attention Network (YOLO11s-APFAN). First, Adaptive Pyramid Focus and Diffusion Network (APFDNet) is introduced to improve feature extraction capability of remote sensing images. Second, Convolution and Attention Fusion Module (CAFM) is integrated to enhance the capacity for adaptive modulation of multi-channel feature representations, thereby improving its precision in localizing small ship targets. Regarding the loss function, Wise-IoU (WIoU) is adopted to enhance the generalization ability of the model across images of varying quality, effectively mitigating the limitations of traditional CIoU in managing partial overlaps and gradient vanishing phenomena. Additionally, the C3K2 module in YOLO11s is enhanced with the PKI approach, utilizing a more efficient parallel convolution combination to better extract key information from images. On the Fog-LEVIR-Ship dataset, the improved algorithm achieved gains of 3.2 % and 1.2 % over the baseline model YOLO11s in mAP@0.5 and mAP@0.5:0.95, respectively. In the generalization experiment on the RS-SSDD dataset, its mAP@0.5 and mAP@0.5:0.95 were improved by 1.9 % and 3.5 %, respectively, compared with YOLO11s.