Huanyu Zhang, Bei Jiang, Ming Chen
This study developed an interpretable machine learning framework based on Random Forest and SHAP analysis for spatiotemporal prediction of pH in public water systems. A systematic comparison of eight machine learning algorithms was conducted, including Linear Regression, Ridge Regression, Lasso Regression, K-Nearest Neighbors, Support Vector Regression, Decision Tree, Gradient Boosting, and Random Forest. Using multi-site monitoring data encompassing dissolved oxygen, specific conductance, temperature, and historical pH values, the Random Forest model achieved competitive and stable performance with test R² of 0.8449, RMSE of 0.0116, and MAE of 0.0066, demonstrating the lowest cross-validation variance among ensemble methods with CV R² of 0.8386 ± 0.0066 among all evaluated models. SHAP interpretability analysis revealed that maximum dissolved oxygen dominates pH prediction with a mean absolute SHAP value of 0.0119, followed by maximum pH and specific conductance indicators. Feature importance validation through a reduced model containing only six key features retained 97.0% of the predictive performance, providing empirical support for monitoring network optimization. Temporal performance evaluation confirmed that the model maintains stable accuracy under diverse water quality conditions with minimal systematic bias. The results confirm that the Random Forest-SHAP framework can provide accurate predictions and mechanistic insights, offering guidance for water quality management and improved understanding of hydrochemical dynamics.