Rabia Korkmaz Tan, Ertuğrul Ordu
The California Bearing Ratio (CBR) is a critical parameter in pavement design and building foundation assessment; however, it requires labor intensive laboratory testing, including a 96 h soaking period. This study evaluated nine machine learning algorithms for predicting CBR from soil index properties: Extra Trees, Support Vector Regression (SVR), Random Forest, Gaussian Process Regression (GPR), CatBoost, LightGBM, XGBoost, Artificial Neural Network (ANN), and ElasticNet. Using 236 soil samples characterized by eight features, we conducted repeated stratified 10-fold cross validation (100 iterations). Extra Trees achieved the highest cross validation R2 of 0.789 ± 0.095 (RMSE = 2.064 ± 0.481, MAE = 1.482 ± 0.294), followed by SVR (R2 = 0.783 ± 0.102, RMSE = 2.090 ± 0.511, MAE = 1.446 ± 0.300) and Random Forest (R2 = 0.777 ± 0.104, RMSE = 2.117 ± 0.460, MAE = 1.518 ± 0.299). The Friedman statistical test confirmed significant performance differences (χ2 = 191.97, p < 10−37), and Nemenyi post hoc analysis identified Extra Trees, SVR, Random Forest, and GPR as statistically equivalent superior groups. SHAP analysis highlighted gravel content (29.0%), maximum dry density (23.8%), and fines content (14.8%), which is consistent with geotechnical principles. Systematic noise injection (20% perturbation) demonstrated model stability, with less than 7% performance degradation at 15% noise. On this heterogeneous compiled dataset, which extends beyond the calibration domain of the empirical equations, all six empirical methods yielded negative R2 (range: −0.803 to −22.639), while all ML models achieved positive R2 (≥0.655 to 0.789). Extra Trees achieved a 3.1× lower RMSE than the best empirical equation, confirming substantially better predictive performance in this out-of-calibration setting. The framework provides a practical five step implementation workflow that may reduce the need for preliminary CBR tests under project specific accuracy thresholds.