Yuzhe Sha, Zhenshan Tan, Xuejin Huo, Rui Liu, Zhanxin Luo, Xianyi Chen
Optical Remote Sensing Image Salient Object Detection (ORSI-SOD) aims to localize visually dominant regions in large-scale remote sensing scenes for applications such as disaster monitoring and urban analysis. Visual saliency fundamentally arises from contrast between a local region and its surrounding context, i.e., the center–surround mechanism. While ORSI-SOD further extends this principle, existing methods still rely on implicit or weak center bias and lack explicit modeling of center–surround spatial contrast, resulting in unstable saliency localization in complex remote sensing scenes. To address this issue, inspired by the human visual system, we propose a saliency detection framework based on SAM2 that explicitly embodies a dynamic center–surround mechanism, termed S2AM. S2AM explicitly reconstructs the saliency localization process by jointly modeling heterogeneous saliency cues, including semantic centers, surround contrast, and boundary constraints in a prompt-free manner. Specifically, we introduce a Saliency-Aware Domain Adapter (SADA) to inject saliency-sensitive activations into generic foundation features, alleviating the weak and implicit center bias inherited from SAM2. Building upon this, a Centroid-Guided Coarse Localization (CGCL) module explicitly predicts semantic centroids and constructs adaptive center–surround contrast structures, enabling robust localization under highly variable object distributions. Finally, a Structure-Constrained Saliency Location Decoder (SCLD) leverages structural cues as spatial constraints to enhance center saliency and suppress surrounding interference. Extensive experiments on the EORSSD, ORSSD, and ORSI-4199 benchmarks demonstrate that S2AM consistently outperforms state-of-the-art methods across multiple evaluation metrics, validating the effectiveness of dynamic center–surround-driven saliency modeling for challenging remote sensing scenarios.