Nora Fino, Ben Haaland, Lesley A Inker, Andrew Levey, Josef Coresh, Ogechi M Adingwupu, Tom Greene
Single-study GFR equations may underestimate cross-study prediction error by ignoring study-level variability, especially in non-CKD populations. Our results largely corroborate the use of pooled datasets and the current form of the CKD-EPI equations. Despite measureable study-level heterogeneity, the pooled CKI-EPI apparach remains a pragmatic and broadly available solution for GFR estimation.
INTRODUCTION: In routine clinical practice, kidney function is assessed by estimating the glomerular filtration rate (GFR) using equations that combine serum creatinine and/or cystatin C with age and sex. These equations are typically created by pooling datasets from multiple research studies and populations into a single dataset. We investigate implications of study and population variation on the relationship of GFR with its predictor variables.
METHODS: We combined data used to develop and validate the 2021 and 2012 The Chronic Kidney Disease Epidemiology Collaboration (CKD-EPI) equations based on creatinine or cystatin C (13,391 individuals in 26 studies). We used random effects modeling to partition the deviations between measured GFR (mGFR) and estimated GFR (eGFR) into four components: predictable variation between studies of CKD vs. non-CKD populations, non-predictable variation between studies, patient-level error in eGFR within studies, and measurement error in mGFR.
RESULTS: Overall, the average eGFR was similar across different levels of the predictor variables between the random effect models and the original pooled CKD-EPI equations. However, the random effect model demonstrated substantial variation between studies, both between and within CKD and non-CKD populations. After accounting for the population type (CKD or non-CKD), approximately 10-40% of the deviations between eGFR and mGFR was attributable to study variation, 20-40% to measurement error in mGFR, and 20-70% to patient level errors in eGFR. Both pooled and random-effects models generalized well across studies, whereas using equations derived from a single study leads to substantial overfitting and larger prediction errors.
CONCLUSIONS: Single-study GFR equations may underestimate cross-study prediction error by ignoring study-level variability, especially in non-CKD populations. Our results largely corroborate the use of pooled datasets and the current form of the CKD-EPI equations. Despite measureable study-level heterogeneity, the pooled CKI-EPI apparach remains a pragmatic and broadly available solution for GFR estimation.