T. Chan, Marcus Yio, Hong S. Wong
This study presents a data-driven analysis of long-term natural carbonation in concrete, using a newly compiled database comprising 1079 mixes and 8194 carbonation depth measurements over 65 years, representing one of the largest natural carbonation datasets assembled to date. Four different tree-based machine learning models (CatBoost, XGBoost, Random Forest and Decision Tree) were evaluated for predicting carbonation rate ( k ), with CatBoost emerging as the most effective. Partial dependence and SHAP analyses were used to quantify the relative importance of 13 features influencing k , including binder composition, mix proportion, curing and exposure conditions. The water-to-calcium oxide (w/CaO) ratio emerges as the most important feature, encapsulating the effects of porosity, carbonatable content, and supplementary cementitious material (SCM) type and content. Higher SCM replacement levels increase carbonation, while aggregate-related features show comparatively less pronounced effects. Carbonation environment is a more important feature than curing condition in predicting k . By combining machine learning with interpretable SHAP analysis, the model independently recognised trends consistent with literature findings. This study offers a comprehensive dataset and robustly selected features to facilitate future machine learning-based predictions of concrete carbonation.