Guangwei Gu, Haonan Shi, Guofang Wang, Juluo Chen
Current predictive models generally demonstrate good overall predictive performance; however, most models suffer from issues such as single-center development, insufficient external validation, and methodological limitations. In the future, more multicenter, large-sample prospective studies should be conducted, and strategies for variable handling and model validation should be optimized to improve the generalizability and clinical translation of predictive models.
BACKGROUND: A laboratory-based frailty index (FI-Lab) summarizes routinely measured deficits, but during critical illness it may reflect acute derangement and underlying vulnerability. Its baseline and dynamic roles require prediction times that precede outcome follow-up.
METHODS: Using MIMIC-IV v3.1, we performed internal temporal validation with separate 24-h and 72-h landmarks. The 24-h cohort comprised patients alive with evaluable 0-24 h FI-Lab; the 72-h cohort comprised those alive with evaluable baseline and 24-72 h FI-Lab. Each outcome was death after the applicable landmark through day 90 after ICU admission. Clinical logistic models were compared with models incorporating baseline FI-Lab, change direction, a four-level phenotype, or continuous baseline and change. We assessed discrimination, calibration, decision curves, and paired bootstrap differences, with sensitivity analyses defined for the revised analysis.
RESULTS: Among 3,376 admissions, 39 deaths occurred by 24 h and 148 by 72 h. The complete-case 24-h cohort comprised 3,186 patients (development: n = 1,890, 559 events; temporal validation: n = 1,296, 339 events). Adding baseline FI-Lab improved validation AUC from 0.774 to 0.808 (paired difference 0.034, 95% CI 0.022-0.047; p < 0.001) and reduced the Brier score from 0.162 to 0.152 (difference -0.011, 95% CI -0.015 to -0.006; p < 0.001). The complete-case 72-h cohort comprised 2,918 patients (development: n = 1,754, 473 events; validation: n = 1,164, 287 events). Baseline FI-Lab improved validation AUC from 0.752 to 0.791. Adding change direction or continuous change yielded AUCs of 0.797 and 0.798, respectively, but neither improvement over baseline FI-Lab was statistically supported; the four-level phenotype had AUC 0.767.
CONCLUSION: Baseline FI-Lab improved internally validated prediction of mortality through day 90 among patients alive and evaluable at 24 h. At 72 h, dynamic change added limited global predictive information, and the phenotype remained exploratory. External validation is required before clinical use.