Jie Xu, Wei Dai, Jason F. Goldberg, Inchi Hu, Chun‐Hung Chen, Palak Shah, Christopher R. deFilippi, Jiayang Sun
The paper evaluates non-interpretable models against interpretable ML models (adaptive logistic regression with interaction terms [aLR], Classification and Regression tree [CART], and conditional inference tree [CIT]) for one-year mortality using balanced SRTR datasets. Predictive validation before the listing policy change in 2018 showed similar AUCs between DNN, aLR, CART, and CIT, with XGBoost having a higher AUC. The interpretable ML models confirm the significance of previously reported risk factors and outperform XGBoost in the post-2018 predictive analysis.
Abstract BACKGROUND Machine learning (ML) models have been used to evaluate one-year post-transplant mortality in donor-recipient pairs. Previous modeling utilizing non-interpretable ML methods (deep neural networks [DNN] and XGBoost) showed modest gains in area under the receiver operating curve (AUC) beyond logistic regression, but suffered a significant drop in predictive AUC applied to the subsequent years’ data and lacked statistical significance in validating previously identified risk factors. METHODS Using balanced SRTR datasets, we evaluated non-interpretable models against interpretable ML models (adaptive logistic regression with interaction terms [aLR], Classification and Regression tree [CART], and conditional inference tree [CIT]) for one-year mortality, including comprehensive clinician-supervised data curation and inclusion of variables describing pre- and post-2018 listing status changes. Models were trained/tested using rolling-window validation across years and further analyzed with repeated ten-fold cross-validation. Interaction terms were obtained via Adaptive Best-Subset Selection (ABESS). RESULTS Predictive validation before the listing policy change in 2018 showed similar AUCs between DNN (0.579), aLR (0.642), CART (0.579), and CIT (0.584), with XGBoost having a higher (0.763) AUC. However, in the post-2018 predictive analysis, aLR outperformed XGBoost (AUC 0.613 vs. 0.586). The interpretable ML models confirm the significance of previously reported risk factors (recipient bilirubin and creatinine) and identify risk factors not previously reported (donor pH, potential recipient distance, and recipient transfusion), as well as clinically relevant interaction terms. CONCLUSIONS Carefully developed interpretable ML models of one-year transplant mortality have similar predictive performance to black-box models, while identifying novel risk factors, and showing improved performance after recent listing policy changes. With appropriate validation and additional data, interpretable ML modeling may allow real-time data-driven donor selection.