Yuhan Liu, Yu Deng
Abstract The rapid advancement of industrial manufacturing has intensified the demand for high-precision and real-time defect detection. However, existing object detection algorithms often face a trade-off between detection accuracy, computational efficiency, and model deployability. To overcome these challenges, this study proposes GSSA-YOLOv8s, a lightweight and efficient detection network derived from YOLOv8s. The model integrates a compact C3K2ScC module to enhance feature extraction while substantially reducing parameters and floating-point operations, thereby minimizing feature redundancy. The incorporation of GhostConv further improves computational efficiency and maintains detection precision under lightweight constraints. Additionally, a shared cross-semantic attention mechanism, combining shared multi-semantic space attention and progressive channel-wise self-attention, establishes a collaborative spatial-channel attention framework to strengthen feature discriminability. An adaptive kernel convolution is also introduced in the detection head to enhance adaptability to geometric deformations. To validate the effectiveness of our proposed GSSA-YOLOv8 model, we conducted comprehensive evaluations on benchmark steel defect datasets. Compared to the YOLOv8s baseline, GSSA-YOLOv8 achieved a notable 2% absolute improvement in mean average precision on the GC10-DET dataset, while simultaneously reducing model parameters by 83.75% and computational demand by 81.4%, as measured in giga floating-point operations. On the NEU-DET dataset, the model maintained its detection accuracy, while realizing a similar reduction of 83.75% in parameters and 81.3% in GFLOPs. These results firmly establish that GSSA-YOLOv8 delivers superior classification accuracy and robustness in the domain of Steel surface defect detection.