Yang Xiao, Jiayi Li, Haijun Dong
Abstract Background Running-related injury (RRI) is a common problem affecting runners’ health and training continuity. Although demographic, training-related, and biomechanical factors have been associated with RRI, the use of multidimensional running data to classify injured and uninjured runners remains insufficiently explored. This study aimed to develop and compare machine-learning models for classifying RRI status and to interpret the best-performing model using Shapley additive explanations. Methods This exploratory cross-sectional study used the University of Calgary Running Injury Clinic database. After data screening, 1,744 runners were included, comprising 1,088 injured and 656 uninjured runners. Candidate predictors included demographic, anthropometric, training-related, running speed, and three-dimensional treadmill running biomechanical features. The dataset was randomly divided into training and independent test sets at a 7:3 ratio using stratified sampling. Feature selection was performed within the training set using least absolute shrinkage and selection operator logistic regression combined with the Boruta algorithm. Five models were developed: logistic regression, support vector machine, random forest, extreme gradient boosting, and LightGBM. Model performance was evaluated using discrimination, classification, and calibration metrics. Results Twenty-two features were retained. In the independent test set, random forest achieved the highest area under the receiver operating characteristic curve among the five models (0.790, 95% confidence interval: 0.747–0.831), with an accuracy of 0.761, sensitivity of 0.905, specificity of 0.523, and F1-score of 0.826. However, a clear training–test performance gap indicated potential overfitting. Shapley additive explanations identified running speed, recreational running level, years of running experience, age, and selected lower-limb biomechanical features as important contributors. In the sensitivity analysis, the biomechanics-only random forest model showed lower performance. Conclusions Machine-learning models, particularly random forest, showed moderate ability to classify RRI status using multidimensional running-related data. These findings provide exploratory evidence for injury status classification, but external validation, prospective studies, and injury subtype-specific modelling are needed before practical application.