Qing Lu, Zhen Cui, Xia Wang
Abstract Objective . Medical image segmentation is persistently hampered by semantic ambiguity, blurred boundaries, and substantial target-scale variation. In clinical settings, these issues coincide as low lesion–background contrast with anatomical clutter, noisy and irregular lesion morphology, and densely distributed small targets with boundary adhesion, thereby constraining conventional methods in semantic representation, structural preservation, and multi-scale modeling. Approach. To address these problems, this paper proposes WSDDANet, a CNN-Transformer hybrid segmentation network that introduces a Dual-Domain Attention (DDA) module and an asymmetric encoder-decoder architecture. DDA adopts a channel-first, spatial-next modeling strategy, and is formed by cascading Wavelet-Enhanced Channel Attention (WECA) and Structure-Enhanced Dynamic Sparse Spatial Attention (SDSSA): WECA leverages discrete wavelet transform and multi-scale channel aggregation to suppress texture noise in the frequency domain and to enhance discriminative channels relevant to lesions; SDSSA employs a dual-level Top-k sparse selection at both region and pixel levels, combined with local context enhancement, to adaptively focus on key structural boundaries. Meanwhile, within the encoder, a standard 3 × 3 convolutional layer is combined with a depthwise-separable operation, whereas the decoder fuses depthwise separable convolution with pyramid convolution, achieving a balance between multi-scale detail reconstruction and computational efficiency. Main Results. We assess our method on three benchmarks: LIDC-IDRI, ISIC2018, and DSB2018, observe consistent gains over competing state-of-the-art models in Dice and mIoU. Significance. These results confirm the efficacy of the proposed approach and demonstrate its consistent performance across the evaluated datasets.