Ningsheng Liao, Yuning Zhang, Zhongliang Yu, Jiangshuai Huang, Mi Zhu, Bo Peng
In the field of computer vision,DEtectionTRansformer (DETR) has received significant recognition for its ability to streamline the design process of object detectors through the concept of set prediction. However, its exceptional performance comes at the cost of a high parameter count and significant computational requirements. Moreover, its ability to detect small objects is compromised, making it less suitable for analyzing high-altitudeUnmannedAerialVehicle (UAV) images. This paper proposes UAV-DETR, a DETR architecture specifically designed for detecting UAV images captured at high altitudes, which achieves a trade-off between parameter count and precision. UAV-DETR is built in two steps: first, inverted residual structures are used to preserve low-dimensional image features, followed by a carefully designed cascaded linear attention mechanism to mitigate parameter redundancy. Through observation and analysis of the attention diffusion issue in the encoder, a cross-channel dynamic sampling mechanism is proposed, which effectively expands the model's receptive field while maintaining accuracy. In addition, the loss function is redesigned by incorporating the Wasserstein distance, which is insensitive to bounding boxes, in order to significantly enhance the convergence speed of the model. Extensive experimental results on two major benchmarks, i.e. VisDrone and UAVDT, validate the simplicity and efficiency of our model. Specifically, on the VisDrone2021 public test set, UAV-DETR exhibits superior performance with only 14 million parameters compared to YOLOv8$_{m}$, reducing the model's parameter count and complexity by 44% and 10% respectively, while achieving a 16.6% improvement in accuracy, without any data augmentation or post-processing procedures.