Milad Jafari, Ehsan Mousavi
Inaccurate and time-consuming construction cost estimation processes during the early stages of projects have long been a critical challenge, prompting researchers to explore alternative costing techniques that leverage historical data and data-driven methodologies. This study conducts a comprehensive analysis of the literature on data-driven construction cost estimation, focusing on research articles published between 2010 and 2024. A novel success ratio metric is introduced to assess the performance of various algorithms in predicting construction costs. Moreover, a detailed statistical analysis of accuracy, sample size, and number of utilized features is provided. The findings reveal that ensemble methods, extreme gradient boosting (XGBoost), case-based reasoning (CBR), and neural networks emerge as the most effective algorithms for construction cost estimation, in descending order of efficiency. A deeper analytical evaluation highlights that selecting an optimal range of 8 to 12 features significantly reduces errors in regression models. Furthermore, the study establishes that using data sets with sample sizes exceeding 200 enhances the robustness and reliability of predictions. Notably, time series-based approaches are found to outperform regression-based methods in terms of accuracy, underscoring their potential for improving cost estimation practices in the construction industry. The findings are particularly helpful for researchers when choosing an analysis technique (or combination of techniques) based on the specific features and characteristics of projects.