Fahim Bordjihene, Salah Eddine Tachi, Hamza Bouguerra, Jazia Arrar
In order to effectively manage water quality, current surface water quality must be closely monitored and evaluated. The Water Quality Index (WQI) is currently used to evaluate surface water quality. However, the calculation of WQI is complicated and time-consuming. In order to optimize the surface water quality assessment of the study area, it is essential to conduct research to develop an efficient and accurate method for calculating the WQI. To this end, eight Machine Learning (ML) regression models were used to predict WQI with reduced input data, thereby reducing the cost of water monitoring. These include XGBoost, AdaBoost, CatBoost, Decision Tree (DT), Random Forest (RF), Multi Linear Regression (MLR), Support Vector Machine (SVM), and Stacking Voting Ensemble Methods (SVEM). The WQI was calculated using water quality data from the Cheffia Dam for 1992–2012 and 2018–2022. The dataset was split into training (70%) and test (30%) groups to develop the models, and feature importance analysis was applied to identify the key parameters to feed the ML models. Model performance was evaluated using statistical metrics, including coefficient of determination ( R 2 ), Mean Absolute Error (MAE), and Root Mean Square Error (RMSE). The results indicate a significant deterioration in water quality during 2018–2022 compared to 1992–2012, with a decrease in excellent water quality and an increase in poor to non-potable water, while NO 2 − , PO 4 3 − , and NH 4 + were the most influential parameters, and all eight models effectively predicted WQI. Among them, the XGBoost model using the third ML03 input combination achieved the highest accuracy ( R 2 = 0.987, RMSE = 0.0181, MAE = 0.0111). In general, these models are useful for predicting the WQI with great accuracy, which improves the assessment and management of surface water quality intended for human consumption.