Nurdaulet Tasmurzayev, Zhanel Baigarayeva, Bibars Amangeldy, Баглан Иманбек, Shugyla Kurmanbek, Gulmira Dikhanbayeva, Gulshat Amirkhanova
Coronary artery disease (CAD) is a leading cause of global mortality, demanding accurate and early risk assessment. While machine learning models offer strong predictive power, their clinical adoption is often hindered by a lack of transparency and reliability. This study aimed to develop and rigorously evaluate a calibrated, interpretable machine learning framework for CAD prediction using 56 routinely collected clinical and demographic variables from the Z-Alizadeh Sani dataset (n = 303). A systematic protocol involving comprehensive preprocessing, class rebalancing using SMOTE, and grid-search hyperparameter tuning was applied to five distinct classifiers. The XGBoost model demonstrated the highest predictive performance, achieving an accuracy of 0.9011, an F1 score of 0.8163, and an Area Under the Receiver Operating Characteristic Curve (AUC) of 0.92. Post hoc interpretability analysis using SHAP (Shapley Additive Explanations) identified HTN, valvular heart disease (VHD), and diabetes mellitus (DM) as the most significant predictors of CAD. Furthermore, calibration analysis confirmed that the mode’s probability estimates are reliable for clinical risk stratification. This work presents a robust framework that combines high predictive accuracy with clinical interpretability, offering a promising tool for early CAD screening and decision support.