Václav Nežerka, Jakub Houlík
Bitumen-aggregate adhesion is a primary determinant of the moisture resistance and durability of asphalt pavements, yet it is still quantified by subjective visual grading, which yields only coarse ordinal categories and varies with the operator. Supervised deep learning can replace it with a continuous, reproducible measurement, but has been constrained by the cost of pixel-level annotation, which requires minutes to hours per photograph. We eliminate this bottleneck by generating the training data synthetically: from 63 manually segmented quarry samples spanning diverse mineralogies, a controlled pipeline composes an effectively unlimited set of pixel-accurate labelled scenes. A U-Net with a ResNet-50 encoder is trained for three-class segmentation (background, bitumen-coated aggregate, stripped aggregate) on these scenes combined with a limited set of manually annotated real photographs. The model attains a mean IoU of 0.83 and a mean Dice of 0.90 on 700 synthetic hold-out images. Preliminary pixel-level validation on 9 previously unseen real photographs yields a mean IoU of 0.76 and a mean Dice of 0.83, with predicted cover on average 1.1 percentage points below manual annotation and estimated 95 % limits of agreement from to points. A separate analysis of 250 synthetic and 147 unique unannotated real photographs returns a domain-classification AUC of 0.998, indicating a substantial residual distribution shift; matched-update ablations show, however, that this shift is dominated by recoverable low-level acquisition differences and that synthetic data support transfer and reduce cover error when combined with real images.