Tianxiang Zhang, Zhaokun Wen, Bo Kong, Kecheng Liu, Yisi Zhang, Peixian Zhuang, Jiangyun Li
Referring Remote Sensing Image Segmentation (RRSIS) is critical for text-guided environmental monitoring, land cover classification, precision agriculture, and urban planning, requiring precise segmentation of objects in remote sensing imagery guided by textual descriptions. This task is uniquely challenging due to the considerable vision-language gap, broad coverage of remote sensing imagery with diverse categories and small targets, and the presence of clustered, unclear targets with blurred edges. To tackle these issues, we proposeSTDNet, a novel framework designed to bridge the vision-language gap, enhance multi-scale feature interaction, and improve fine-grained object differentiation. Specifically,STDNetintroduces (1) the Spatial Multi-Scale Correlation (SMSC) for improved vision-language feature alignment, (2) the Target-Background TwinStream Decoder (T-BTD) for precise distinction between targets and non-targets, and (3) the Dual-Modal Object Learning Strategy (D-MOLS) for robust multimodal feature reconstruction. Extensive experiments on the benchmark datasets RefSegRS and RRSIS-D demonstrate thatSTDNetachieves state-of-the-art performance, effectively dealing with the core challenges of RRSIS with enhanced precision and robustness. Consequently, it is envisaged that the proposedSTDNetmodel will be an advantage in the RRSIS task. Datasets and codes are available athttps://github.com/wzk913ysq/STDNet.