科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Frontiers in medicine2026-01-01· Missing data

Development and internal validation of machine learning models for gouty arthritis classification using routine clinical variables: a retrospective study.

Weiwei Ma, Weikang Sun, Zhiyong Hu, Weidong Liang, Huanan Li

一句话结论 · In one sentence

A machine-learning model based on routinely available hospital variables, particularly random forest, showed favorable internal discrimination for GA classification and outperformed conventional logistic regression. These findings support its use as a preliminary screening-oriented classification aid rather than a stand-alone diagnostic tool. External validation remains necessary before broader clinical deployment.

原始摘要(英文原文)· Original abstract
BACKGROUND: Gouty arthritis(GA) is a common crystal-induced inflammatory arthropathy influenced by inflammatory, metabolic, renal, hematologic, demographic, and lifestyle-related factors. Because serum uric acid alone may not fully characterize GA in routine clinical settings, integrative models based on routinely available variables may improve disease classification. METHODS: This single-center retrospective study included 7,383 eligible adult participants from the electronic medical record database of the Affiliated Hospital of Jiangxi University of Chinese Medicine between September 1, 2020 and February 1, 2026. GA cases were retrospectively reconstructed according to the 2015 American College of Rheumatology/European League Against Rheumatism classification framework, and non-gout comparators were drawn from a contemporaneous heterogeneous hospital population. Missing predictor values were handled using multiple imputation by chained equations after stratified training/test splitting. Least absolute shrinkage and selection operator regression was used for feature selection. Four supervised learning algorithms were developed and internally evaluated using a 7:3 hold-out design, with cross-validation-based tuning within the training set. Model performance was assessed using discrimination, calibration, decision curve analysis, and confusion matrices. RESULTS: Among the four models, random forest showed the best overall performance. In the training set, it achieved an area under the receiver operating characteristic curve of 0.8995, accuracy of 0.7062, sensitivity of 0.9623, positive predictive value of 0.6094, and negative predictive value of 0.9419. In the test set, it remained the best-performing model, with an area under the receiver operating characteristic curve of 0.8686, accuracy of 0.6980, sensitivity of 0.9442, positive predictive value of 0.6048, and negative predictive value of 0.9163. Support vector machine ranked second, whereas k-nearest neighbors and logistic regression showed weaker discrimination. CONCLUSIONS: A machine-learning model based on routinely available hospital variables, particularly random forest, showed favorable internal discrimination for GA classification and outperformed conventional logistic regression. These findings support its use as a preliminary screening-oriented classification aid rather than a stand-alone diagnostic tool. External validation remains necessary before broader clinical deployment.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Development and internal validation of machine learning models for gouty arthritis classification using routine clinical variables: a retrospective study. — 科研速览 Science Skim