Lina Gao, Haikun Chen, Yonggang Zhang, Yulong Huang
Although salient object detection (SOD) methods inspired by human attention mechanisms have received increasing attention for their superior performance, they suffer from a trade-off between efficacy and efficiency, due to the high performance generally coming at the expense of large parameter sizes and high computational costs. To address this issue, we propose a Collaborative Prior-Enhanced RGB-D Salient Object Detection Network (CPENet) for Intelligent IoT Perception Devices. Specifically, we propose a dual-stream architecture utilizing MobileViT as the backbone to tackle the challenge of capturing long-range dependencies among multi-modal features with reduced model complexity. Then, a Scale Information Enhancement Module (SIEM) is proposed to effectively enhance and fuse multi-scale features. To adaptively aggregate multi-modal features, we propose a Prior-Driven Modality Aggregation (PDMA) that integrates the coarse saliency maps (Prior-Glance) of model into cross-modal fusion. Furthermore, we propose an extremely simple shared decoder to both predict the saliency maps of fused features for deep supervision and generate Prior-Glance in generation modules. Finally, considering the deficiency of boundary representation, unlike other methods which append costly edge enhancement module, a prior-driven region loss is designed that improves the capacities of boundary representation and generalization of the model under the guidance of Prior-Region. Extensive experiments demonstrate that the proposed CPENet outperforms most state-of-the-art methods on seven datasets with fewer parameters (15.7M) and lower computational complexity (21.9GFLOPs). Codes and results are available at https://github.com/blossom-lv/CPENet.