科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Cureus2026-07-01

The Accuracy of the I2 Statistic for Detecting Baseline Imbalances and Selection Bias in Randomized Controlled Trials: A Simulation Study.

Steffen Mickenautsch, Veerasamy Yengopal

一句话结论 · In one sentence

The I2 test was significantly more accurate for identifying baseline imbalance and selection bias detection, and equally accurate for ruling out baseline imbalance compared with statistical significance testing. However, the I2 test was significantly less accurate in detecting selection bias absence. The I2 test's bias category scores were largely negatively associated with p-values and tended to overestimate bias severity. Negative I2 test results appear reliable for ruling out baseline imbalances and selection bias in RCTs, while positive I2 test results are highly accurate to detect baseline imbalances but should be interpreted cautiously when assessing selection bias.

原始摘要(英文原文)· Original abstract
AIM: To assess the accuracy of the I2 statistic, applied as a trial-adjusted, simulated comparator trials (SCTs) based I2 test for detecting baseline imbalances and selection bias in randomized controlled trials (RCTs). METHODS: Fifty non-biased and 50 biased trials set at 40% selection bias severity were simulated in MS Excel (Microsoft Corporation, Redmond, Washington, USA). All trials were tested using statistical significance testing and the I2 test. Null hypotheses were tested that the sensitivities and specificities of both tests did not statistically significantly differ and that the selection bias category scores, established by the I2 test, do not statistically significantly correlate with the p-values from statistical significance testing. McNemar test and Spearman's rank correlation were used. The risk of the I2 test overestimating selection bias severity was also examined. RESULTS: I2 test demonstrated higher sensitivity than significance testing for detecting both baseline imbalance and selection bias (p < 0.0001). Specificity for baseline imbalance detection was identical at 100% for both methods; 95% CI: 90%-100%. For selection bias, specificity was greater with significance testing than with the I² test (p = 0.0005). A significant negative, large (0.5 ≤ |r|) correlation emerged between I² scores and significance test p-values: Spearman's r = -0.61, p < 0.0001. All null hypotheses, with the exception of the test specificities for baseline imbalance detection, were rejected. The I2 test caused 26% of 'true positive' trials to be over-classified by one severity category. CONCLUSION: The I2 test was significantly more accurate for identifying baseline imbalance and selection bias detection, and equally accurate for ruling out baseline imbalance compared with statistical significance testing. However, the I2 test was significantly less accurate in detecting selection bias absence. The I2 test's bias category scores were largely negatively associated with p-values and tended to overestimate bias severity. Negative I2 test results appear reliable for ruling out baseline imbalances and selection bias in RCTs, while positive I2 test results are highly accurate to detect baseline imbalances but should be interpreted cautiously when assessing selection bias.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

The Accuracy of the I2 Statistic for Detecting Baseline Imbalances and Selection Bias in Randomized Controlled Trials: A Simulation Study. — 科研速览 Science Skim