Lulu Liu, Rui Zhu, 郭辰嘉, Yushuai Zhang, Zhibo Shi, Yaru Li, Le Gao
Unmanned aerial vehicle (UAV) imagery at low altitude is increasingly used for traffic oversight, infrastructure inspection, and emergency awareness, yet bird's-eye views exhibit pronounced scale diversity, minute object footprints, and cluttered backgrounds. Default YOLO-style single-stage detectors on a P3-P5 pyramid can therefore suffer coupled cross-scale misalignment and within-scale receptive-field mismatch, increasing misses and unstable calibration on tiny instances. Built on the Ultralytics YOLO11 macro graph and training recipe for strict comparability, RBD-YOLO addresses these bottlenecks with three coordinated methods: (R) within-scale adaptive receptive-field reweighting, (B) learnable cross-scale weighted fusion, and (D) scale-space-task joint dynamic calibration. Under one unified protocol, with VisDrone-DET validation as the primary benchmark and UAVDT as an auxiliary cross-dataset check, RBD-YOLO achieves 49.6% precision, 39.4% recall, and 40.1% mAP@0.5 on VisDrone, outperforming the same-series YOLOv11n baseline (44.4%, 33.2%, 33.1%) and remaining ahead of newly added larger-scale YOLO-family comparisons in mAP@0.5. On UAVDT, it reaches 98.2% mAP@0.5 and 95.4% recall, indicating competitive cross-dataset behavior on a nearly saturated vehicle-centric benchmark. Ablations attribute decomposable gains to methods B and R, a dominant increment from widening fusion bandwidth, and further refinement from method D once cross-scale capacity is sufficient; at 640 × 640 the full model uses about 3.47 M parameters and 16.4 GFLOPs, yielding an interpretable accuracy-complexity trade-off for VisDrone-like aerial small-object detection.