Bo Liu, Juan Shi, Hamed Karimian, Fei Wang, Shengnan Shi, Haochen Wang, Hui Liu, Yuehua Chen
• A lightweight, highly accurate model is proposed for underwater object detection. • Our newly designed neck reduces the parameters by nearly 10%. • The proposed model detects a wider range of small underwater objects. • An Efficient Multi-Scale Attention (EMA) module is added to the backbone’s feature pyramid structure to enhance detection precision. Accurate object recognition in the marine environment is critical for protecting biodiversity and ensuring the sustainable exploitation of ocean resources. While state-of-the-art techniques based on deep learning have demonstrated impressive accuracy, their high computational requirements prevent real-time deployment, particularly on resource-constrained platforms such as AUVs and ROVs. To overcome this limitation, we present RED-YOLO, a lightweight yet accurate model tailored for underwater vision. Firstly, we introduced an Efficient Multi-scale Attention (EMA) module into the backbone to strengthen multi-resolution feature extraction. In addition, the neck structure incorporates a Lightweight-RepGFPN design to simplify feature fusion while reducing redundant parameters. Finally, we replaced DCNV2 with DCNV4, which employs a scale-, spatial, and task-aware attention mechanism to facilitate target localization. Our model outperforms other models with just 2.8M parameters and 9.7 GFLOPs, achieving 87.3% mAP on our comprehensive test dataset. We believe that RED-YOLO provides a promising solution for underwater object detection by striking a balance between computational efficiency and detection accuracy.