Ju-Hyeon Noh, Junyoung Jang, J. H. Jo, Hee-Deok Yang
Crack detection and analysis are essential for maintaining the stability and longevity of infrastructure; however, traditional manual inspections or simple image processing techniques are inefficient. To address this, automated crack segmentation using deep learning is being actively researched. This study proposes a hybrid model combining U-Net and a Vision Transformer to enhance the accuracy of crack segmentation. The proposed model is based on U-Net’s encoder–decoder architecture and integrates a Convolutional Neural Network (CNN), which is strong in local feature extraction, with a Vision Transformer, which excels at capturing global features and long-range dependencies, to effectively learn complex crack patterns. Experimental results on the CrackSeg9k dataset show that the proposed model achieves a mean Intersection over Union (mIoU) of 0.7184, demonstrating superior segmentation performance compared to other models like the conventional U-Net and Attention U-Net. This indicates that the proposed hybrid approach successfully leverages both local and global features, proving its effectiveness in segmenting complex and irregular crack patterns.