Yang Li (7082), Xing Wu, Chengliang Wang, Hao Wu, Peng Wang, Hongqian Wang
Scribble supervision provides a cost-effective alternative to full annotation for medical image segmentation by offering sparse structural cues, but its imprecision often leads to semantic confusion and inaccurate segmentation. Recently, the general purpose Segment Anything Model(SAM), which exhibits strong structural completion ability, has been introduced into scribble-supervised settings to reconstruct complete object structures from sparse prompts. However, due to the domain gap between natural and medical images, the quality of SAM-generated pseudo-labels still needs improvement. Moreover, overly simplistic applications of SAM fail to fully exploit its structural modeling capability, leaving semantic confusion largely unresolved. To address the above challenges, we propose a three-stage closed-loop iterative framework, termed SAM-SP, for scribble-supervised medical image segmentation. Specifically, in the first stage, a contour-based prompt selection strategy is employed to extract structurally constrained information from scribble annotations, guiding SAM to generate initial pseudo-labels. In the second stage, which serves as the core optimization stage, a category prototype-based semantic alignment and cross-scale semantic fusion mechanism is introduced to effectively alleviate semantic confusion and transfer the structural prior knowledge of SAM to the segmentation network. In the third stage, an uncertainty-guided pseudo-label fine-tuning strategy is further adopted to adaptively refine SAM’s knowledge blind spots in medical images, thereby narrowing the domain gap. Extensive experiments demonstrate that the proposed method achieves significant performance improvements across multiple medical segmentation benchmarks, with Dice improvements ranging from 2.7% to 39.1% and a maximum reduction of 26.3 in HD95, validating its effectiveness in mitigating semantic confusion and enhancing weakly supervised segmentation performance.