Vosco Pereira, Oseko Yutaka, Hidekazu Fukai
Identifying road surface cracks by semantic segmentation is a difficult problem. This is because segmentation typically detects objects by area, whereas cracks are string-like. Conventional loss functions such as Binary Cross-Entropy (BCE), Dice, and IoU often fail to capture the fine, elongated features of cracks, as they rely on pixel-level, area-based overlap, leading to suboptimal performance. To address this, we investigate one of the skeleton-based losses, the Centerline Dice (clDice) loss, which emphasizes the preservation of tubular structures via soft skeletonization. We improve road crack segmentation by combining clDice with conventional loss functions, systematically evaluating its role by varying the weight parameter and skeletonization iterations. Experiments are conducted on the EdmCrack600 and CrackForest datasets using two segmentation models: a customized CNN-based U-Net++ and a transformer-based SegFormer. Performance is evaluated using the Dice coefficient, IoU, clDice, and Hausdorff Distance. Results show that combining clDice and IoU loss with customized U-Net++ achieves superior performance. Compared to a standard BCE baseline, it improves the Dice coefficient by 4.9 and 2.8 percentage points on EdmCrack600 and CrackForest and improves the clDice score by 3.9 and 1.7 percentage points. These results highlight improved segmentation of thin, linear cracks, supporting practical advancements in road monitoring and segmentation of linear structures.