科研速览 · Science Skim继续刷下去 · Keep skimming →
◇ medRxiv2026-08-26· epidemiology

Quantifying the Optimism of Naive Cross-Validation for Binary Outcome Prediction with Repeated-Measures Predictors: A Simulation Study and Clinical Illustration

J. Hagan

原始摘要(英文原文)· Original abstract
Background. Cross-validation (CV) is widely used to estimate the discriminative performance of clinical prediction models, but applied at the observation level to repeated-measures data it can overestimate performance. A common but under-characterized structure in neonatal, critical care, and wearable-sensor research pairs continuous predictors measured repeatedly within subjects with a binary subject-level event, evaluated by area under the receiver operating characteristic curve (AUROC). The magnitude of this optimism remains uncharacterized against true out-of-sample performance. Methods. Three strategies for estimating subject-level AUROC from ridge logistic regression were compared: naive observation-level 10-fold CV, subject-level 10-fold CV, and leave-one-cluster-out (LOCO) CV. They were applied to daily oxygenation measures and retinopathy of prematurity (ROP) outcomes in 101 extremely low birth weight infants, and to a factorial simulation of 162 conditions varying cluster count (20 to 150), intraclass correlation (ICC, 0.1 to 0.5), lag-1 autocorrelation (0.2 to 0.8), and prevalence (10% to 35%). An independent sample provided a true-performance reference, decomposing optimism into a partitioning component and a residual discrepancy. A paired 24-condition arm held latent signal variance constant across ICC. Results. In the clinical dataset, naive CV optimism was +0.078 AUROC units for severe ROP and +0.031 for any ROP. In the simulation, total optimism relative to true performance averaged +0.213, comprising a partitioning component (mean +0.154, range +0.071 to +0.257) and a residual internal-validation discrepancy (mean +0.059). The partitioning component increased with higher ICC and autocorrelation, fewer clusters, and lower event rates, whereas the residual discrepancy did not rise with dependence. Under signal control, the rise in the partitioning component across ICC was largely retained (+0.041 versus +0.052) whereas the decline in the residual was attenuated by 73%. Subject-level 10-fold CV closely agreed with LOCO throughout (mean absolute deviation 0.004 in simulation). Conclusions. Naive observation-level CV meaningfully overestimates discriminative performance; subject-level partitioning eliminates the partitioning component and closely agrees with LOCO, although a residual discrepancy from true performance remains at small cluster counts. Subject-level partitioning should be considered essential rather than optional, at every level of nested resampling.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Quantifying the Optimism of Naive Cross-Validation for Binary Outcome Prediction with Repeated-Measures Predictors: A Simulation Study and Clinical Illustration — 科研速览 Science Skim