Chuanwei Zhao, Honglin Li, Yane Yang, Ying Yang, Xinhua Wu
The fixed eight-predictor XGBoost model provided useful cross-hospital risk discrimination for CA-AKI, but its incremental advantage over same-variable logistic regression was limited and external calibration was imperfect. Prospective multicenter validation, standardized predictor timing, and local calibration assessment are required before clinical implementation.
BACKGROUND: Contrast-associated acute kidney injury (CA-AKI) is a clinically relevant complication after coronary angiography in patients with acute coronary syndrome (ACS). We aimed to develop and geographically validate a compact, interpretable machine-learning model and to determine whether the nonlinear algorithm provided incremental value beyond conventional regression using the same predictors.
METHODS: This retrospective two-center prediction study included patients with ACS undergoing coronary angiography. The derivation cohort was obtained from Dali, and the Baoshan cohort was reserved for locked geographical external validation. The primary XGBoost model used eight fixed predictors: age, blood urea nitrogen, baseline serum creatinine, fibrinogen, lymphocyte percentage, renin-angiotensin system inhibitor use, uric acid, and glycated hemoglobin A1c. Internal validation used nested five-fold cross-validation after establishment of the predictor set with fold-specific preprocessing. Performance was assessed using discrimination, calibration, prediction error, decision curve analysis, exploratory risk stratification, sensitivity analyses, and SHAP-based interpretation.
RESULTS: The derivation cohort included 2,677 patients, of whom 302 developed CA-AKI, and the external validation cohort included 1,539 patients, of whom 243 developed CA-AKI. The fixed eight-predictor XGBoost model achieved an internal out-of-fold area under the receiver operating characteristic curve (AUROC) of 0.876 and an external AUROC of 0.819. In external validation, the area under the precision-recall curve was 0.581, the Brier score was 0.096, the calibration slope was 0.747, and the observed-to-expected event ratio was 1.059. XGBoost substantially outperformed models based on baseline serum creatinine, estimated glomerular filtration rate, or basic clinical variables. However, its improvement over logistic regression using the same eight predictors was modest. Risk categories preserved clear risk ordering across hospitals, although lower-risk patients were underpredicted. Reduced-predictor models retained similar external performance, whereas complete-case analyses were limited by substantially smaller sample sizes.
CONCLUSIONS: The fixed eight-predictor XGBoost model provided useful cross-hospital risk discrimination for CA-AKI, but its incremental advantage over same-variable logistic regression was limited and external calibration was imperfect. Prospective multicenter validation, standardized predictor timing, and local calibration assessment are required before clinical implementation.