科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Remote Sensing2025-11-18· Computer science

MSFFNet: Multimodal Spatial–Frequency Fusion Network for RGB-DSM Remote Sensing Image Segmentation

Yuanjie Zhi, Yuhang Wang, Fan Zhang, Mingyang Ma, Shaohui Mei

原始摘要(英文原文)· Original abstract
Remote sensing image segmentation is essential for resource planning and disaster monitoring. Although RGB-based methods are widely adopted, they often exhibit suboptimal performance in distinguishing objects with similar color and texture characteristics. The fusion of height information from Digital Surface Models (DSM) aids in the discrimination of these challenging objects. However, existing CNN- and pooling-based fusion methods tend to lose edge details as network depth increases, resulting in blurred segmentation boundaries. To address this issue, a Multimodal Spatial–Frequency Fusion Network (MSFFNet) is proposed to effectively enhance edge details by fusing high-level frequency and spatial features. Specifically, a Hybrid Branch Fusion Module (HBFM) is proposed, in which the wavelet transform branch decomposes features into sub-components, effectively isolating edge and structural information from other textures. Such a process in the frequency domain prevents edge details from being lost or diluted during fusion, thereby preserving boundary clarity in segmentation. Additionally, a Multi-Scale Contextual Attention Module (MSCAM) is proposed to capture multi-scale contextual information for enhancing spatial feature representation, while adjusting both spatial and channel-wise attention mechanisms to improve detail and accuracy. Experiments over benchmark Vaihingen and Potsdam datasets demonstrate that the proposed approach can clearly enhance edge delineation while improving segmentation accuracy.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

MSFFNet: Multimodal Spatial–Frequency Fusion Network for RGB-DSM Remote Sensing Image Segmentation — 科研速览 Science Skim