科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ International journal of medical informatics2026-09-10

Scalable multiple imputation for national-scale healthcare registries and EHR pipelines.

Hugo Morvan, Jonas Agholme, Björn Eliasson, Katarina Olofsson, Ludger Grote, Fredrik Iredahl, Oleg Sysoev

一句话结论 · In one sentence

bigMICE can make principled multiple imputation feasible for national-scale healthcare registries, thereby reducing reliance on complete-case analysis in large observational studies and facilitating more robust clinical and epidemiological pipelines.

原始摘要(英文原文)· Original abstract
BACKGROUND: Missing data remain a major challenge in healthcare quality registries and electronic health records (EHRs). Multiple Imputation by Chained Equations (MICE) is widely used for handling missingness in clinical research in a statistically rigorous fashion, but conventional implementations of MICE often become computationally infeasible for modern registry-scale datasets because of memory and/or runtime limitations. This substantially limits clinical researchers, who often do not have access to powerful computational facilities needed to run MICE on registry-scale data. OBJECTIVE: To develop and evaluate a scalable framework for multiple imputation in large healthcare registries that can be run on a standard personal computer. METHODS: We developed bigMICE, an open-source framework that limits memory usage through the dynamic exchange of data between memory and storage while distributing computations across the threads of the local machine. We evaluated the framework using a large Swedish National Diabetes Registry (NDR) dataset under varying sample sizes, number of variables and missingness levels. Performance metrics included runtime, memory consumption and imputation quality. Using NDR and other registries, bigMICE was benchmarked against conventional, non-distributed MICE and the deep learning-based imputation method MIWAE. RESULTS: Compared with the non-distributed MICE approach, bigMICE reduced peak memory usage for the largest NDR subsets from 40 GB to 11 GB and computational time by ∼80%, enabling registry-scale MICE imputations on a standard PC. Both bigMICE and MIWAE consumed less than 16 GB of RAM, while bigMICE required up to ∼50 times less computational time. Imputation quality of bigMICE remained stable up to 99% missingness, suggesting that highly incomplete variables can still aid imputation of large datasets. CONCLUSIONS: bigMICE can make principled multiple imputation feasible for national-scale healthcare registries, thereby reducing reliance on complete-case analysis in large observational studies and facilitating more robust clinical and epidemiological pipelines.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Scalable multiple imputation for national-scale healthcare registries and EHR pipelines. — 科研速览 Science Skim