Mohammad Bahram, Hossein khosravi
ABSTRACT Accurate prediction of the compressive strength of geopolymer concrete is essential for designing sustainable and cost-effective construction materials. This study evaluates and compares the performance and accuracy of four machine learning models data, and R = 0.123 and RMSE = 16.37 for the testing data, indicating poor generalization capability. In contrast, Jafari et al. achieved better performance using more advanced algorithms, with R—XGBoost, LightGBM, CatBoost, and Support Vector Regression (SVR)—for predicting compressive strength based on 162 geopolymer concrete mix designs. Each sample includes 16 input features, such as curing temperature, chemical composition of sodium silicate and fly ash, superplasticizer content, water, fine and coarse aggregates, and oxides (Fe₂O₃, Al₂O₃, CaO, Na₂O, SiO₂), with compressive strength as the target variable. All data processing and analysis were performed using Python software.To assess the relative performance of the developed models, the results were compared with previous studies. Loukog et al. employed machine learning models to predict the compressive strength of geopolymer concrete and reported a correlation coefficient of R = 0.768 and RMSE = 8.764 for the training = 0.89 and RMSE = 4.83 for the training set, and R = 0.826 and RMSE = 5.96 for the testing set. Compared with these studies, the models developed in the present research—particularly the XGBoost model—demonstrated substantially higher accuracy. In the testing set, XGBoost achieved a correlation coefficient of 0.86 and an RMSE of 5.59, outperforming both previous works. Moreover, other Boosting-based models (CatBoost and LightGBM) also exhibited competitive results, performing similarly to or better than previous studies. This comparison highlights that the adoption of Boosting algorithms can significantly enhance the prediction accuracy and stability of compressive strength estimation in geopolymer concrete.In addition to these performance advantages, the application of these predictive models results in considerable savings in laboratory time and costs, while significantly reducing human errors that may occur during traditional experimental procedures. The use of these algorithms facilitates faster and more optimized mix design, minimizing the risk of errors associated with manual testing.