Zhaisheng Ding, Ruichao Hou, Yunzhe Men, Shengyang Luan, Yanyu Liu, Kangjian He, Shidong Xie
An effective knowledge learning strategy combined with a lightweight network architecture is crucial for the practical deployment of multi-modal image fusion. While existing methods have made significant progress in the visual perception of fused results, their model complexity and generalization capabilities still require further optimization. In this paper, we propose a novel passive-active distillation learning framework for multi-modal image fusion, termed DSKFuse, which integrates the Dynamic Sparse Transformer and the latent Kolmogorov-Arnold Network (KAN). Specifically, we design an efficient fusion architecture trained via a two-stage knowledge distillation strategy, seamlessly integrating passive and active learning methodologies. In the first stage, passive distillation learning enhances the fusion network by extracting valuable knowledge from complex fusion models. In the second stage, an active knowledge distillation approach is implemented, enabling the model to autonomously capture discriminative features from source images, thereby improving the robustness and generalization of DSKFuse. Extensive experiments demonstrate that the proposed method achieves state-of-the-art performance in both image fusion and downstream tasks, including detection and segmentation. The code will be released at https://github.com/DZSYUNNAN/DSKFuse .