Jinqing Zheng, Zhiyong Feng, Meng Xing, Yong Su, Weilong Peng, Yiming Zhang
Oriented Vehicle Detection (OVD) in drone imagery plays a fundamental role in a wide range of downstream applications, owing to its capacity to provide accurate localization and corresponding categorization of vehicle instances in complex scenes. Existing OVD methods for drone-based scenes are limited by their coarse granularity, while real-world applications often require fine-grained detection—for example, distinguishing between van trucks and tank trucks is critical for traffic regulation and road safety. To address this gap, we introduce Fine-Grained Oriented Vehicle Detection (FGOVD) task for the first time in drone-based scenarios. To facilitate this task, we develop a Semi-Automatic Labeling approach for Fine-Grained Oriented Object Detection, named SAL-FGOOD, and construct a truly fine-grained multi-scene dataset, Drone Fine-Grained Oriented Vehicle Detection (DFGOVD) dataset. SALFGOOD extends the existing labeling paradigm by incorporating a category-adjustment branch, reducing manual workload by 94.7% for our DFGOVD dataset and 77.7% for the public FAIR1M-v2 dataset. Building upon SAL-FGOOD, we construct DFGOVD, the first large-scale and comprehensive FGOVD dataset for drone-based scenes. The dataset comprises 33,669 images and 816,239 vehicle instances spanning 53 categories and captures subtle distinctions in vehicle appearance, size, and functionality. We benchmark state-of-the-art oriented object detection methods on the DFGOVD dataset, exposing its inherent complexity and challenge. The dataset is available upon request from https://github.com/JinqingZhengTju/DFGOVD.