Sandra Ekelund, Julie Rosenskjold, Thomas Kronborg, Stine Hangaard
BACKGROUND: COPD is a leading cause of morbidity and mortality worldwide. However, 80% of individuals with COPD are undiagnosed, leading to accelerated disease progression and increased healthcare costs. Different screening and case-finding tools have been suggested, yet only a few have explored the use of machine learning (ML) in this context. Thus, the aim of this study was to develop an ML-based classification model, utilizing only primary care data, to identify people with undiagnosed COPD. METHODS: /FVC < 0.7. Four different ML models were trained and evaluated, and performance was assessed by sensitivity, precision, and area under the receiver operating characteristic curve (AUROC). For the best-performing model, based on AUROC, differences in feature characteristics between the true positive group and the false negative group were explored. RESULTS: A total of 11,265 participants were included, 449 with undiagnosed COPD and 10,816 with non-COPD. The best-performing model was the logistic regression model with an AUROC of 0.84 (95% CI 0.81-0.87) and a sensitivity of 0.72 (95% CI: 0.64-0.79), but a low precision of 0.12 (95% CI: 0.11-0.14), indicating a high number of false positive classifications by the model. The false negative classifications were markedly younger, had fewer pack-years, and had a greater proportion of never-smokers compared with the true positive group. CONCLUSIONS: An ML-based classification model, exclusively using information available in primary care, could identify individuals with undiagnosed COPD. However, further development to increase model precision and external validation of the model is required to ensure generalizability before implementation in clinical settings.