Ali Akbari, Mojtaba Rahimi
This study addresses the limitations of traditional methods, such as Archie’s equation and conventional log interpretation, by applying nine machine learning (ML) techniques (LR, MLR, GPR, SVM, SVR, RF, ANN, DT, and kNN) for predicting water saturation (Sw). The input parameters included RXO, RD, GR, CNL, SP, and depth, and the contribution of each parameter to model accuracy was evaluated. To enhance performance, the PSO algorithm was employed for model optimization. In addition, Gaussian elimination was used for outlier detection and removal, allowing performance comparison before and after data cleaning. Model robustness was further assessed under different data availability conditions by varying the training/testing ratios from 10% to 90%, an aspect that has been rarely explored in previous research. The results indicated that the DT model achieved the best performance with an R2 of 0.9437 on test data, 0.9775 on training data, and an RMSE of 0.019 on test data. The kNN model showed an R2 of 0.9288 on test data and 1.0 on training data, with an RMSE of 0.0 on test data and 0.026 on training data. The SVM model achieved an R2 of 0.9469 on test data, 0.9890 on training data, and an RMSE of 0.01 on test data and 0.02 on training data. ML methods provide more accurate Sw estimation than conventional techniques, offering reliable, cost-effective alternatives to core sampling and lab tests in the oil and gas industry. The proposed workflow can also be readily applied to other heterogeneous carbonate reservoirs, highlighting its potential for broader practical implementation in reservoir characterization.