Alexander Pate, G. Martin, Richard D. Riley
Abstract Introduction The heuristic shrinkage factor of Van Houwelingen and Le Cessie (๐ ๐๐ป ) is a commonly used closed-form solution to adjust for overfitting in unpenalised logistic regression models for risk prediction. It is also the basis of widely-adopted minimum sample size criteria for developing clinical prediction models. However, current evidence is lacking regarding the bias of ๐ ๐๐ป compared to the optimal shrinkage factor (๐ ๐๐๐ก ). Here, we examine this issue and also assess the bias of an alternative bootstrap-derived shrinkage factor (๐ ๐๐๐๐ก ). Methods We undertook two simulation studies. The first examined the bias of ๐ ๐๐ป and ๐ ๐๐๐๐ก as estimators of ๐ ๐๐๐ก across a range of different scenarios defined by ๐ถ ๐๐๐ , the C-statistic of the model developed in a population sized dataset. The second examined the bias of ๐ ๐๐๐ก when using development sample sizes targeting a shrinkage of 0.9, based on a sample size calculation defined by ๐ ๐๐ป itself (๐ ๐๐๐๐๐๐๐๐ ) or by an adapted simulation-based approach (๐ ๐ ๐๐ ). Results For high C-statistics, ๐ ๐๐ป overestimated ๐ ๐๐๐ก , whereas for low C-statistics ๐ ๐๐ป underestimates ๐ ๐๐๐ก . For example, across scenarios when 0.8โค๐ถ ๐๐๐ <0.85, the 95-percentile range in the bias was (0.005,0.387), compared to (โ0.580,โ0.007) across scenarios when 0.6โค๐ถ ๐๐๐ <0.65. The magnitude of bias increased as ๐ถ ๐๐๐ tended to either 0.5 or 1. As sample size increased and ๐ ๐๐๐ก โ1, the magnitude of the bias in either direction reduced. ๐ ๐๐๐๐ก was less biased than ๐ ๐๐ป , with a median magnitude of bias across all scenarios of 0.007, compared to 0.032 for ๐ ๐๐ป . Developing models on datasets of size ๐ ๐ ๐๐ gave ๐๐๐๐(๐ ๐๐๐ก ) closer to 0.9 (mean magnitude of bias across all scenarios 0.004) than ๐ ๐๐๐๐๐๐๐๐ (mean magnitude of bias 0.041). Conclusions ๐ ๐๐ป is often a poor estimator of the optimal global shrinkage factor. If global shrinkage is needed, we recommend using the bootstrap shrinkage estimate. The bootstrap estimate shows minimal bias in most scenarios, though in small samples the variability is large so provides no guarantees to address overfitting in a single dataset. A sample size calculation based on simulation is often preferable over formula dependent on targeting ๐ ๐๐ป .