Fengqian Sun, Deqiang Cheng, Ping Zheng, Tianshu Song, Liangliang Chen, Kou Qiqi
Recent studies have highlighted the importance of contextual information for small object detection. However, existing methods rely solely on visual features and lack additional semantic guidance, which limits their ability to model key scene-level context in semantically rich, globally complex environments and to suppress irrelevant local context in densely cluttered scenes. These limitations hinder their effectiveness in Unmanned Aerial Vehicle (UAV) and similar complex scenes. To address these challenges, we propose TGCADNet (Text-Guided Context-Aware Detection Network). TGCADNet is a small object detection framework that leverages the CLIP (Contrastive Language– Image Pretraining) model’s global semantic understanding and image-text alignment capabilities for enhancing context-aware detection. TGCADNet mainly consists of Text-Guided Scene-level Context-Aware (TG-SCA) and Text-Guided Local-Context Filtering (TG-LCF). Specifically, TG-SCA uses CLIP-generated text features to guide the model in accurately extracting key scene-level context from globally complex environments. Meanwhile, TG-LCF performs interactive computation between text and image features to filter high-quality local context, thereby reducing the impact of dense and cluttered local regions in UAV scenes. We validate the effectiveness of TGCADNet on the VisDrone, UAVDT, and AI-TOD-v2 datasets. Compared to the baseline, TGCADNet achieves an improvement of 1.8 in mAP@50 and 1.3 in mAP@50:95 on the VisDrone dataset. On the UAVDT and AI-TOD-v2 datasets, TGCADNet observes improvements of 2.5 and 2.3 in mAP@50, respectively. Furthermore, TGCADNet surpasses recent SOTA methods in both accuracy and efficiency, demonstrating its effectiveness in detecting small objects in UAV and similar remote sensing scenes.