Zhuangzhuang Du, Xiangdang Tong, Yinfeng Hao, Xinhui Zhou, Lei Zhang, Junfeng Tian, Chengquan Zhou
Accurate identification of fish feeding intensity is essential for intelligent feeding control and cost reduction in aquaculture. However, under complex aquaculture conditions, water surface reflections, glare interference, and high background noise often degrade the reliability of single-modal recognition methods. Although multimodal models can alleviate these limitations, they usually suffer from large parameter scales, cross-modal alignment difficulties, and limited feasibility for edge deployment. To address these issues, this paper proposes a Heterogeneous Multi-level Knowledge Distillation framework. A multi-stage audio-visual fusion network is constructed, incorporating a reliability-aware gating mechanism, a bidirectional residual interaction fusion backbone, and a robust adaptive decision fusion module to enhance cross-modal complementary representations under complex operating conditions. Meanwhile, a multi-level distillation strategy across the decision, feature, and relationship levels is designed to transfer class boundary information, key intermediate representations, and cross-modal collaborative structures from the teacher model. Furthermore, to suppress erroneous knowledge transfer caused by high-noise samples and modality-misaligned samples, an uncertainty-aware defense mechanism based on teacher prediction entropy is introduced. This mechanism dynamically adjusts the distillation intensity, enabling lightweight student networks to efficiently inherit deep latent knowledge from high-capacity teacher models. Experiments on the public AV-FFIA dataset demonstrate that the constructed student model achieves a recognition accuracy of 97.47% while reducing the number of parameters by approximately 60%, with performance approaching that of the high-capacity teacher model. The proposed method provides a feasible solution for lightweight deployment of multimodal feeding intensity recognition in complex aquaculture environments.