Zezhao Meng, Yubo Zheng, yuhang jin, Mingbo Wu, Haiqiang Zuo
Abstract: Surface defect segmentation plays a crucial role in the fields of intelligent manufacturing and quality control. However, restricted by objective factors in industrial scenarios such as complex background textures, drastic variations in defect scale, and similarity between classes, existing segmentation methods often fail to achieve ideal accuracy. Traditional convolutional neural networks (CNNs) are limited by local receptive fields, making it difficult to capture long-range dependencies; meanwhile, while Vision Transformers excel at global modeling, their representation of low-level pixel details is insufficient, easily leading to missed detections or blurred segmentation of fine features such as small-scale defects and defect edges. Addressing the aforementioned practical challenges in industrial scenarios and the inherent defects of existing CNN and Transformer methods, this paper proposes HDSNet, a dual-branch hybrid network fusing CNN and Transformer. Specifically, it achieves precise surface defect segmentation by forming a complementary relationship between the local detail extraction advantages of CNNs and the global semantic modeling advantages of Transformers. To improve network performance, HDSNet incorporates three core innovations. First, a local–global fusion module is proposed to combine the advantages of convolutional local detail extraction and self-attention-based global modeling in a parallel manner. This module interleaves and concatenates global features from the two branches along the channel dimension, generates cross-channel attention weights through one-dimensional convolution, and introduces learnable parameters to adaptively balance the proportion of features from dual branches. It realizes the efficient fusion of local and global features and enhances the model’s capability to fully perceive large-scale defects. Second, a large Kernel attention downsampling module is introduced, utilizing large receptive field attention to replace traditional pooling operations, which maximizes the preservation of the spatial structural information of defects while reducing feature dimensions and computational overhead. Finally, a grouped dual attention module is designed, which, through grouping strategies and channel-spatial dual-path attention, effectively distinguishes background textures from defect features, significantly suppressing noise interference in complex backgrounds. HDSNet achieves the best segmentation performance on three different public industrial surface defect datasets. HDSNet outperforms current mainstream methods in terms of both segmentation accuracy and robustness, demonstrating good prospects for industrial application.