Linfeng Jiang, Peidong Zhan, Ting Bai
Traffic sign detection is a vital component of intelligent transportation systems. However, in real-world driving scenarios, challenges such as illumination variations, occlusions, and low resolution of small objects can significantly reduce detection accuracy. To overcome these challenges, we propose YOLO-MAFF, a traffic sign detection network that integrates a multi-scale attention mechanism and adaptive feature fusion. Firstly, a backbone network incorporating a multi-scale channel attention mechanism is designed. By integrating multi-scale contextual information with channel attention, efficient feature extraction and representation learning are facilitated. Secondly, a pyramid network based on adaptive feature fusion is developed to learn spatial attention maps. By fusing feature maps at various scales and emphasizing or suppressing region-specific features, the network can alleviate inconsistencies in feature representations. Finally, a small object detection layer is designed to preserve shallow-level detail information in the feature maps, enabling the network to detect small traffic signs. In the experimental section, YOLO-MAFF is evaluated on four datasets, i.e., TT100K, CCTSDB2021, CURE-TSD, and COCO. The experimental results show that YOLO-MAFF exhibits superior performance in traffic sign detection tasks. Compared to the baseline YOLOv8s, our method improves the mAP by 4.8% on TT100k (reaching 90.2%), 2.8% on CCTSDB2021 (reaching 86.0%), 2.7% (reaching 53.7%) on CURE-TSD, and 2.0% (reaching 72.5%) on the COCO dataset. The source code is available athttps://github.com/lfjiang-cn/yolo-maff