Mohammadreza Saberironaghi, Jing Ren, Alireza Saberironaghi
Automated pixel-level detection of steel surface defects is a critical challenge in manufacturing quality control, complicated by the variation in defect size and shape, low contrast with background textures, and the diversity of defect patterns. This paper proposes ID-MSNet, an enhanced version of the UNet3+ architecture, designed specifically for the segmentation of three common steel surface defect types: inclusions, patches, and scratches. The proposed architecture introduces three targeted modifications: (1) a multi-scale feature learning module (MSFLM) in the encoder that uses dilated convolutions at multiple rates to capture contextual features across different scales, combined with DropBlock regularization and batch normalization to improve generalization; (2) an improved down-sampling (IDS) module that replaces standard max-pooling with learnable strided convolutions fused via 1 × 1 convolution, preserving richer feature representations; and (3) a convolutional block attention module (CBAM) integrated into the skip connections to selectively focus the model on spatially and channel-wise relevant defect regions. Experiments on the publicly available SD-saliency-900 dataset demonstrate that ID-MSNet achieved an 86.19% mIoU, outperforming all compared state-of-the-art segmentation models while using only 6.7 million parameters—approximately 75% fewer than the original UNet3+. These results establish ID-MSNet as a strong and efficient baseline for steel surface defect segmentation, with potential applicability to automated quality inspection in broader manufacturing contexts.