Shaozhong Zhang, Pengfei Shao, Haidong Zhong, Yaohui Wu, Chenjie Du
Accurate biomass estimation is crucial for evaluating production capacity and resource efficiency in shrimp aquaculture. Traditional monitoring and measurement approaches typically rely on single-modality data, such as sonar, water quality sensors, or images. However, these data sources are often susceptible to environmental disturbances and fail to capture the complexity of aquatic ecosystems, resulting in unstable estimation accuracy. This study proposes a novel multimodal deep learning approach to improve biomass estimation. It leverages the complementary information provided by video surveillance and water quality sensor data to address the challenge of sensitivity to environmental noise in unimodal prediction. To tackle the challenges of fusing these heterogeneous modalities, we introduce a Contrastive Learning and Attention-based Cross-modal Fusion (CACF) model. The model employs a contrastive learning framework for cross-modal alignment and utilizes a self-attention mechanism to dynamically weight features, ensuring robust data fusion. Comprehensive experiments conducted on a real-world shrimp biomass dataset demonstrate that CACF significantly outperforms state-of-the-art models, improving both the accuracy and reliability of biomass estimation. This work highlights the potential of multimodal data integration in aquaculture and lays the foundation for the future development of real-time estimation systems and intelligent feeding technologies to optimize shrimp farming practices. • A multimodal deep learning model integrates video and sensor data for accurate shrimp biomass estimation. • Cross-modal contrastive learning ensures semantic and temporal alignment of features. • A self-attention mechanism adaptively fuses modalities based on feature relevance. • Real-world experiments show significant accuracy gains over existing baseline methods.