游双红, Jinquan Li, Lingyun Yao, Chongxi Yan, Zheng Wu
Automated peach picking in complex orchard environments faces challenges such as fruit overlap, occlusion by leaves and branches, and low efficiency in continuous localization, which require detection algorithms to achieve both high accuracy and real-time performance. To address these issues, this study proposes a lightweight and high-precision peach detection model named Peach-YOLO based on an improved YOLOv8n framework. First, a Receptive-Field Attention Convolution (RFAConv) module is introduced into the C2f structure of the backbone network to enhance feature representation in salient regions and suppress background interference. Second, a Convolution and Attention Fusion Module (CAFM) is integrated to further strengthen the extraction of key fruit features. In the neck network, a Coordinate Attention-guided high-level screening feature fusion pyramid network (CA-HSFPN) is adopted, which significantly reduces computational complexity while improving semantic representation. Furthermore, the Shape-IoU loss function is introduced to replace the traditional CIoU loss, achieving more accurate bounding box regression through geometric alignment that accounts for object shape. Experimental results on a custom peach dataset show that Peach-YOLO, with a compact model size of only 5.0 MB, achieves a real-time inference speed of 115.7 FPS, an mAP@0.5 of 82.2%, a precision of 78.9%, and a recall of 76.4%. Compared with the baseline YOLOv8n, the mAP, precision, and recall are improved by 3.0, 3.6, and 4.8 percentage points, respectively. Compared with current mainstream detection models, Peach-YOLO demonstrates superior performance in both accuracy and efficiency, providing a lightweight, high-precision, and real-time visual detection solution for automated fruit picking systems.