Haoxue Zhang, Gang Xie, Linjuan Li, Chenhao Chang, Xinlin Xie, Jinchang Ren, Heng Li
Semantic segmentation of multimodal remote sensing images plays a crucial role in geospatial analysis. However, existing methods often struggle to effectively integrate spatial–spectral and geometric–geographic information, capture multiscale contextual features, and preserve the fine-grained boundaries of geo-objects. To address these challenges, we propose an interaction-guided multimodal fusion network with elevation constraint, referred to as IGECNet. The proposed network comprises four key components. First, a multimodal correlation interaction mechanism adaptively fuses features from high-resolution remote sensing images and digital surface models (DSMs) using multiscale dilated convolutions and cross-attention mechanisms, thereby enhancing complementary relationships while preserving modality-specific characteristics. Second, a channel reorder enhancement module hierarchically decomposes channel-wise features based on entropy and global average pooling values, improving spatial–spectral discriminability. Third, a frequency-guided context decoder refines segmentation outputs by leveraging global semantic context via Fourier transform and fine boundary details through local feature fusion. In addition, a new loss function, jointly constrained by semantic, boundary, and elevation consistency, enforces elevation consistency from DSMs gradients and boundary alignment. Extensive experiments on the ISPRS Vaihingen, ISPRS Potsdam, and DroneDeploy Segmentation benchmarks demonstrate that IGECNet achieves state-of-the-art performance. The source code will be publicly available upon publication.