Wajeeha Iftikhar, Muhammad Yaseen, Gohar Rahman, Muhammad Nauman, Umar Farooq Khattak
Diabetes can cause a lot of serious health problems, and its early detection is very important. This study proposes a hybrid machine learning framework that enhances diabetes prediction accuracy and interpretability by combining the Synthetic Minority Over-Sampling Technique (SMOTE) and Shapley Additive Explanations (SHAP), examining five machine learning models—Logistic Regression (LR), Support Vector Machine (SVM), Decision Tree (DT), Random Forest (RF), and XGBoost—on two datasets. The results showed that RF and XGBoost achieved the highest predictive performance. SMOTE improved class balance and model robustness, while SHAP provided transparent explanations of key predictors such as glucose, BMI, and age. The proposed approach demonstrates that the combination of SMOTE and SHAP enhances both the reliability and the interpretability of models for practical diabetes prediction.