Wentao Wu, Fanghua Hong, Xiao Wang, Chenglong Li, Jin Tang
Vehicle detection is a fundamental perception task in intelligent transportation systems and plays a crucial role in enabling reliable traffic perception and analysis. Existing vehicle detectors are typically obtained by training conventional object detection models (e.g., YOLO, RCNN, and DETR series) on vehicle images based on pre-trained backbone networks (e.g., ResNet and ViT). Although some studies introduce large-scale foundation models to improve detection performance, these models are not specifically designed for vehicle-centric scenarios and therefore tend to yield sub-optimal results in complex traffic environments. Moreover, most existing methods heavily rely on visual features and pay limited attention to the alignment between vehicle semantic information and visual representations. In this paper, we propose a novel vehicle detection paradigm, termed VFM-Det, which integrates a pre-trained vehicle foundation model (VehicleMAE) with a large language model (T5) to achieve semantically enhanced vehicle detection for intelligent transportation scenarios. Specifically, the proposed method follows a region proposal-based detection framework and employs VehicleMAE to enhance the features of each proposal. More importantly, we introduce a novel VAtt2Vec module to predict the vehicle semantic attributes corresponding to each proposal and transform them into feature vectors, which further enhance visual features through contrastive learning. Extensive experiments on three vehicle detection benchmark datasets thoroughly proved the effectiveness of our vehicle detector. Specifically, our model improves the baseline approach by +6.0%, +8.4% on the$AP_{0.5}$,$AP_{0.75}$metrics, respectively, on the Cityscapes dataset. The source code of this work will be released athttps://github.com/Event-AHU/VFM-Det