Hanyun Li, Linsong Xiao, Lihua Cao, Sai Yao, Minghao Wang, Yi Li
Machine vision-based anti-drone detection systems enable long-range, cost-effective target monitoring in complex environments. However, small drones typically occupy only a few pixels in captured images. Existing detectors suffer from semantic loss and insufficient fusion during feature extraction and cross-scale interaction, resulting in limited detection accuracy. To address these challenges, this paper proposes Diffusion Focusing Former (DFFormer), a detection framework specifically designed for small target identification. The framework employs a backbone network to extract multi-layer features, which are enhanced through an Advanced Feature Processing Layer (AFPL) to strengthen semantic representation. A Feature Scaling Layer (FSL) then organically fuses shallow and high-level information before encoder processing, preserving fine-grained cues while minimizing computational overhead. Subsequently, the Multi-Scale Focusing Diffusion Network (MSFDN) processes scaled features for cross-scale interaction and progressive fusion. The Focusing Fusion Module (FFM) injects comprehensive contextual information into each scale throughout this process. Experimental results on three anti-drone datasets (DUT-Anti-UAV, Bird-UAV, and Anti-UAV (Inf)) demonstrate that DFFormer consistently outperforms existing state-of-the-art methods across multiple evaluation metrics. Generalization validation on the VisDrone2019 aerial dataset further confirms the method’s applicability to diverse scenarios and configurations.