Cyril Jaksic, PhDThomas Perneger, Christophe Combescure
Background: When low-power analyses yield statistically significant results, they likely overestimate the true effect. Although sample estimates are symmetrically distributed around the true value, those that are by chance very high are more likely to achieve statistical significance. The bias induced by the significance filter increases as power decreases. Here we sought to quantify the estimation bias associated with low power and to contrast it with the type M error, which assesses the same phenomenon from a different perspective. Methods: We used simulations to quantify estimation bias in relation to power among statistically significant results. We computed the type M error, relative bias (ratio of the estimated mean differences and the true value), and proportions of results with various levels of over- and under-estimation. Results: For a medium effect size (Cohen's d of 0.5), overestimation of the mean difference was moderate at high power (≥80%): relative bias was <1.13, about 65% of estimates were roughly accurate (between 0.75 and 1.25 of the true value), and sign errors were virtually absent. In contrast, at low power (<30%), overestimation was strong (relative bias >1.78), and almost no estimates were roughly accurate. Sign errors became noticeably prevalent only at very low levels of power (<10%). In all situations, the relative bias had a lower magnitude than the type M error. Conclusion: Low-power statistically significant results may consist entirely of magnitude errors, sign errors, and type 1 errors with high risk of strong overestimation (double effect). Readers should beware positive results from low-power analyses.