科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ IEEE transactions on neural networks and learning systems2026-09-18

DIFF-MF: A Difference-Driven Channel-Spatial State Space Model for Multimodal Image Fusion.

Yiming Sun, Zifan Ye, Qinghua Hu, Pengfei Zhu

原始摘要(英文原文)· Original abstract
Multimodal image fusion aims to integrate complementary information from multiple source images to produce high-quality fused images with enriched content. Although existing approaches based on state space models (SSMs) have achieved satisfactory performance with high computational efficiency, they tend to either over-prioritize infrared intensity at the cost of visible details, or conversely, preserve visible structure while diminishing thermal target salience. To overcome these challenges, we propose DIFF-MF, a novel difference-driven channel-spatial SSM for multimodal image fusion. Our approach leverages feature discrepancy maps between modalities to guide feature extraction, followed by a fusion process across both channel and spatial dimensions. In the channel dimension, a channel-exchange module enhances channel-wise interaction through cross-attention dual state space modeling, enabling adaptive feature reweighting. In the spatial dimension, a spatial-exchange module employs cross-modal state space scanning to achieve comprehensive spatial fusion. By efficiently capturing cross-modal discrepancy features and integrating them in a well-balanced manner, DIFF-MF effectively fuses complementary multimodal information. Experimental results on the driving scenarios and low-altitude unmanned aerial vehicle (UAV) datasets demonstrate that our method outperforms existing approaches in both visual quality and quantitative evaluation.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

DIFF-MF: A Difference-Driven Channel-Spatial State Space Model for Multimodal Image Fusion. — 科研速览 Science Skim