Jiali Han, Yannan Deng, Guangwei Zhao, Haotian Shen, Chaoshi Zhang, Yuxuan Xu, Jinhe Zhang, Xixi Zhao, Sixiang Liang, Tong Guo
RF showed the strongest repeated AUPRC among the evaluated algorithms and was selected as the final model, but the small number of events, very limited case detection at the locked threshold and absence of external validation preclude clinical deployment.
BACKGROUND: Non-suicidal self-injury (NSSI) is clinically important among people with mood disorders. We developed an interpretable machine-learning model for NSSI during one-year follow-up and evaluated its internal robustness.
METHODS: The cohort comprised 912 inpatients with mood disorders (MD) and 72 candidate predictors. We collected demographic information, disease characteristics and laboratory indicators of MD patients who were hospitalized throughout 2022, and followed them up one year after discharge to collect their NSSI incidence within one year. Least Absolute Shrinkage and Selection Operator (LASSO) regression was utilized to identify predictors associated with NSSI features. Subsequently, multiple machine learning algorithms were constructed, and the optimal model was selected based on its area under the curve (AUC), F1-score, and accuracy. Finally, Shapley Additive Explanations (SHAP) were applied to the optimal model to rank feature importance and interpret the directional impact of each selected predictor on NSSI risk.
RESULTS: Fifty-two participants (5.70%) experienced NSSI. Across 25 repeated internal evaluations, random forest (RF) had the highest mean AUPRC (0.456; SD 0.116) and the most fold-level AUPRC wins (8/25) and was therefore designated the final model. The prediction of RF achieved a ROC-AUC of 0.861 (95% CI 0.738-0.951), AUPRC of 0.285 (95% CI 0.172-0.535) and Brier score of 0.070; at the locked threshold, recall was 0.091 and F1 was 0.133. RF SHAP importance was highest for age, previous NSSI, FT4, current NSSI and estradiol.
CONCLUSION: RF showed the strongest repeated AUPRC among the evaluated algorithms and was selected as the final model, but the small number of events, very limited case detection at the locked threshold and absence of external validation preclude clinical deployment.