Marion Soto, Ulrich Schimmack
The replication crisis heightened interest in methods for assessing the credibility of published research. One approach is to estimate the average power of original studies based on observed data. Critics have challenged this approach as an “ontological error”, as a poor predictor of replication outcomes, and as too imprecise to be useful. This article aims to address these critics. We respond by clarifying that using observed data to estimate true power is a standard inferential practice and does not constitute an ontological error. The goal of average power estimation is not to predict the outcome of future replication studies, but the hypothetical outcome if original researchers had to replicate their studies with new samples. Lastly, we demonstrate that even with substantial uncertainty, average power estimates remain diagnostically informative, especially under selection for statistical significance. Using a z-curve analysis of terror management research, we illustrate that a corpus of significant findings does not, by itself, imply evidential value under severe selection for significance. We conclude that, despite limitations, average power estimation remains a valid tool for evaluating published research.