科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Research square2026-07-27

Machine learning approaches for tuberculosis prevalence and risk factor association among high-risk groups in Kigali health facilities.

Theophilla Igihozo, Fiacre Rugamba Rugero, Emma Marie Umutoniwase, Emelie Ndoli, Emmanuel Niyonshuti, Rwubaka Eugene, Emmanuel Sibomana, Melissa Uwase, Dieudonne Kayiranga, Michael Mugisha, Celestin Twizere, Isambi Sailon Mbalawata, Davila-Roman Victor, Adams Wilcox, Randi Foraker, Isaac Gomina, Michelle Schneider, David Tumusiime, Emmy Mugisha

一句话结论 · In one sentence

This first application of EMR-driven ML for identifying TB risk factor associations in Rwanda shows that ensemble models, combined with rigorous leakage control, achieve promising performance among high-risk populations. These tools could support national TB surveillance and screening to enhance early detection and resource allocation.

原始摘要(英文原文)· Original abstract
BACKGROUND: Tuberculosis (TB) remains a major global health challenge, disproportionately affecting vulnerable populations in resource-limited settings. In Rwanda, the burden is high among high-risk groups including prisoners, mining workers, people living with HIV, and healthcare workers. Despite well-established national surveillance infrastructure, published evidence on applying machine learning (ML) to routinely collected electronic medical records (EMR) for identifying factors associated with TB in these populations remains limited. This study develops and compares multiple ML models to identify factors associated with TB among high-risk groups using national surveillance data from Kigali health facilities, while emphasizing careful feature engineering to avoid information leakage. METHODS: We conducted a retrospective, cross-sectional analysis of 2,254 high-risk patient EMR from Rwanda Biomedical Center (RBC) surveillance systems, with TB status defined by GeneXpert results. Rigorous data cleaning and preprocessing were performed, including exclusion of records with indeterminate outcomes, removal of diagnostic and operational variables that could introduce information leakage, imputation of missing values, normalization of numerical features, and one-hot encoding of categorical variables. Multiple supervised ML models were developed and evaluated using a stratified train-test split, including logistic regression (LR), balanced random forest (BRF), gradient boosting methods (XGBoost, LightGBM, and GBM), support vector machines (SVM) with a radial basis function kernel, and an EasyEnsemble classifier. Model performance was assessed using accuracy, precision, recall, F1-score, and area under the receiver operating characteristic (ROC) curve (AUC). RESULTS: After excluding indeterminate GeneXpert outcomes, 358 cases were retained for evaluation. Ensemble-based models showed consistently strong and stable performance, with BRF achieving the highest discrimination (AUC = 0.86), while GBM, SVM, and EasyEnsemble classifiers showed comparable performance across accuracy, precision, recall, and F1-score. Feature-importance analysis identified clinically meaningful factors, associated with TB, including site of disease, age, TB-related body mass index, HIV status, and diabetes status, indicating that the models relied on plausible biological and demographic risk factors rather than diagnostic artifacts. CONCLUSIONS: This first application of EMR-driven ML for identifying TB risk factor associations in Rwanda shows that ensemble models, combined with rigorous leakage control, achieve promising performance among high-risk populations. These tools could support national TB surveillance and screening to enhance early detection and resource allocation.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Machine learning approaches for tuberculosis prevalence and risk factor association among high-risk groups in Kigali health facilities. — 科研速览 Science Skim