Seth S. Leopold
Fret not, and please don’t turn the page. Despite the somewhat intimidating title of this month’s Spotlight article [7], you don’t have to be a biostatistician—or even good at mental math—to enjoy it. And it’s not just for foot and ankle specialists, either. The paper, this Spotlight essay, and the Take 5 interview that follows with its first author are for anyone who cares whether the findings we read in clinical journals can be safely applied to patients. They’re also for anyone interested in conducting clinical trials, or understanding how trial design translates to an ability to answer surgically relevant questions. Taken together, that should be all of us. Let’s back up a minute. About 20 years ago, John Ioannidis—an epidemiologist, methodologist, and physician at Stanford University—published a paper that has been cited over 14,660 times in which he claimed that “most published research findings are false” [3]. It’s a fascinating essay that gets deep into the arithmetical weeds, but its main, well-substantiated contentions included: •The smaller the studies conducted in a scientific field, the less likely the research findings are to be true; •The smaller the effect sizes in a scientific field, the less likely the research findings are to be true; •The greater the number and the lesser the selection of tested relationships in a scientific field, the less likely the research findings are to be true; and, intriguingly, •The hotter a scientific field (with more scientific teams involved), the less likely the research findings are to be true. Although his paper was published in a journal you may not follow [3], it’s freely available online, and you should give it a look. His reasoning in it is sound, and it’s the kind of cautionary tale we need more of. Then, about 8 years ago, Dr. Ioannidis [4] published a potential mitigating solution to the problem he surfaced in that earlier paper [3]—or, I should say, he amplified a solution endorsed by “72 methodologists” (of whom he was one) [1]. That group suggested that in medical specialties where authors tend to default to a p value threshold of < 0.05 to define statistical significance for “new discoveries”—which is essentially all medical and surgical specialties—that threshold should be lowered to < 0.005, and that p values between 0.05 and 0.005 should be referred to as “suggestive” rather than “significant.” They claimed that doing so would “improve the reproducibility of scientific research in many fields.” Although these suggestions were published in a leading journal JAMA [4]—they strike me as iffy or worse, for many reasons. One reason I believe these ideas are iffy was explored in this month’s Spotlight paper by Yoshiharu Shimozono MD, PhD (Fig. 1), and his team at Kyoto University in Japan: Moving the “routine” p value down by 10x would predictably result in the need for much larger clinical trials—nearly 70% larger, in fact, according to Dr. Shimozono’s paper—if we want those trials to be adequately powered to detect differences among the treatments being considered [7]. Since it’s unlikely clinician-scientists will suddenly increase the size of all the studies they perform (if that were easy, it’d have been done long ago), indiscriminately making p values 1/10th of what they’d been would leave us awash in a sea of underpowered studies, which doesn’t seem so helpful. And we saw this in the study by Dr. Shimozono’s team: More than half of the randomized clinical trials they evaluated that claimed statistical differences would have been reclassified as no-difference studies or as only “suggestive” of a difference had the p value threshold of 0.005 been used. We need to be thoughtful about how we use p values if we’re going to use them at all. We don’t want studies that erroneously claim differences, but we also don’t want to miss differences that are in fact present because our studies are too small to detect them. I find this topic thorny for many other reasons, as well, as I’ve been saying for quite some time now [6]. (With very little to show for it, I should add; I guess my consolation here is that other than his 15,000 citations, Dr. Ioannidis’s paper [3] hasn’t made much of a dent, either.) Specifically, I find the uncritical use of any single p value threshold (whether 0.05, 0.005, or any other) to be, well, just silly. There’s no a priori basis in logic for doing so. Critiques of arbitrary uses of p values go back almost as far as the description of p values themselves [2], and using thresholds so unselectively is not mirrored in real life by any other well-adapted human decision-making process of which I’m aware. Happily, I believe that clinician-scientists have many better options to choose from than the blind embrace of any single p value threshold, my two favorites being Bayesian inference-testing methods and simply moving the selected p value threshold up or down depending on what one is studying. For example, one might select a more stringent p value threshold if the topic may result in recommending a treatment that carries avoidable risk, like discretionary surgery or a medication with potentially serious side effects. On the flip side, in studies about educational interventions or other topics that don’t involve exposing patients to potential harms, more-relaxed p value thresholds (like 0.1 or even higher) may be fine, provided that the conclusions are presented modestly. This type of “Bayes Through the Back Door” approach reflects what clinicians do in practice every day anyway. For example, clinicians often use histological analysis of frozen sections to look for a prosthetic joint infection (PJI). Since published “cutoffs” for this test range from as low as 2 polymorphonuclear cells per high-powered field (pmn/hpf) to 15 pmn/hpf, the savvy clinician might move the threshold up and down based on a pretest estimation (best guess) of the likelihood of infection. If the surgeon believes that infection is probably present, then 5 pmn/hpf might suffice as confirmatory evidence, but if the clinical presentation suggests PJI is not so likely, perhaps 15 pmn/hpf will be needed to convince the surgeon that infection is, in fact, present. Likewise, a thoughtful reader might articulate a “pre-read probability” of a new study’s contention being true, based on prior clinical experience or on the available evidence published, before consuming the study in question. In other words, if a contention appears to be more of a long-shot based on what is already known, it should take a lower (stricter) p value to alleviate a reader’s skepticism (Fig. 2).Fig. 1: Yoshiharu Shimozono MD, PhDFig. 2: The “Bayes Through the Back Door” approach for choosing the most appropriate p value when designing or consuming clinical research [6].A thoughtful clinician-scientist could save his or her readers a step by setting the p values accordingly when designing his or her study at the outset. Although I made this suggestion in print here 8 years ago [6], and at numerous meetings for years before that while wearing my “editor-in-chief hat,” only a handful of papers have taken me up on it. That’s a missed opportunity. And, of course, all this statistical talk misses the ball altogether if authors don’t report on effect sizes, and if journal editors don’t require them (we do, here at CORR®). Patients don’t perceive p values, they only perceive effect sizes. Reporting on effect sizes using metrics like minimum clinically important differences (MCIDs), patient acceptable symptom states, and substantial clinical benefits—and presenting the uncertainty around effect size estimates—is more helpful to the surgeons who read journals and the patients whom they treat. For now and for the foreseeable future, though, the p value questions being debated in this month’s Spotlight article are vitally important if we’re to be able to trust what we read. This topic sounds methods-heavy, but it’s not beyond the reach of any curious reader, and it really does deserve your attention. So don’t miss the plain-language Take 5 interview that follows below with Yoshiharu Shimozono MD, PhD, first author of, “What Would Be the Effect of Lowering the Threshold of Statistical Significance From 0.05 to 0.005 in Foot and Ankle Randomized Controlled Trials?” [7]. Take 5 Interview with Yoshiharu Shimozono MD, PhD, first author of “What Would Be the Effect of Lowering the Threshold of Statistical Significance From 0.05 to 0.005 in Foot and Ankle Randomized Controlled Trials?” Seth S. Leopold MD:Congratulations on this thought-provoking and important article [7]. You’ve been immersed in this topic now for awhile. I think we agree that unselective application of any single p value threshold to all clinical research makes little sense. If you could choose from the menu of “better” alternatives Dr. Ioannidis once offered [4]—which include not using any threshold, moving the threshold up and down depending on what is being studied, abandoning p values altogether, using Bayesian or other approaches to inference testing, reporting only effect sizes and confidence intervals, among others—what would you recommend that clinician-scientists (and the journals that publish their work) choose and why? What would be the tradeoffs of your proposed approach? Yoshiharu Shimozono MD, PhD: After analyzing 85 foot and ankle RCTs, I believe we need flexible, risk-adjusted p value thresholds rather than abandoning p values entirely, which would be impractical given their entrenchment in medical education and regulatory frameworks. The threshold should reflect multiple factors: the irreversibility of the intervention, the availability of alternatives to it, the costs to patients and healthcare systems, and the quality of the available evidence. For irreversible procedures with limited evidence base—like novel surgical techniques, procedures that eliminate joint motion, or expensive biologics with unclear mechanisms—we should require p < 0.005 or stricter. Our analysis [7] showed that 52% of “positive” surgical RCTs would be interpreted differently if stricter thresholds were used. Conversely, for interventions that are reversible, low cost, and low burden—such as physical therapy protocols, activity modifications, or orthotic adjustments—the traditional p < 0.05 threshold remains reasonable, though not sacrosanct. However, p values alone are insufficient. We must mandate reporting of effect sizes, confidence intervals around those effect size estimates, and clarity on whether any differences exceed MCIDs. Too often in foot and ankle surgery, we celebrate statistically significant but clinically meaningless improvements, like 3 points on a 100-point scale, a difference far too small for patients to perceive. Additionally, we should incorporate informal Bayesian reasoning. For example, if a study were to claim that arthroscopic debridement helped patients with ankle osteoarthritis at p = 0.04, we should pause to consider biological plausibility and prior evidence. One marginally significant study shouldn't overturn our understanding of cartilage biology. The main challenge of making these kinds of changes would be increased complexity. Researchers would need to justify their chosen threshold based on intervention invasiveness, reversibility, cost, and existing evidence quality, which requires a sophisticated understanding of both statistics and clinical context. There's also the risk of inappropriate manipulation, where researchers might select thresholds to achieve “significance.” To prevent this, journals should require threshold specification at the protocol stage, and deviations should trigger rejection. The sample size implications are substantial. Our analysis [7] showed that achieving adequate power with p < 0.005 requires approximately 70% larger samples. For a typical hallux valgus study, this means 170 patients instead of 100—challenging but achievable. For more uncommon conditions, however, adequately powered RCTs would become essentially impossible. Instead of being deterred by these challenges, however, we must view them as opportunities that will ultimately strengthen our field by forcing more thoughtful research design and multicenter collaboration. Professional societies could assist in these efforts, too, by developing nuanced guidelines based on intervention characteristics rather than broad categories—considering factors like reversibility, cost, existing evidence quality, and availability of alternatives. Dr. Leopold:One thing that I do like about a proposed shift from 0.05 to 0.005 is that it highlights the concept that our clinical knowledge is critically contingent on the sizes of the studies that we read. It’s fair to say that larger studies would give us greater confidence in what we think we know. But large studies are expensive; for example, one trial about venous thromboembolism seeks to enroll 25,000 patients, and the cost is estimated to be USD 17 million [5]. That’s to address just one research theme, in the setting of just two operations (hip and knee replacement). It's hard to imagine getting anywhere close to that scale for all of the many important questions we all want answered. What then should we do as a specialty? Dr. Shimozono: We need to prioritize our research ruthlessly based on disease burden, potential for harm, and evidence gaps. Priority should go to common conditions where substantial practice variation indicates genuine clinical uncertainty—where treatment decisions currently depend more on surgeon training and preference than robust evidence. These warrant well-powered multicenter trials despite the cost. For conditions where treatment variation has minimal impact on outcomes, or where natural history is favorable regardless of intervention, we should acknowledge this rather than conducting underpowered RCTs that create false certainty. We also need systematic data collection through registries or collaborative networks. While challenging to implement, standardized outcome tracking across multiple centers could provide the sample sizes needed to answer important clinical questions. The infrastructure investment would be substantial, but considering the costs of treatments performed without strong evidence, it represents good value. Additionally, alternative study designs—adaptive trials, platform trials, or pragmatic trials embedded in clinical practice—could help maximize efficiency with limited resources. Finally, we need intellectual honesty about evidence limitations. When treating conditions where RCTs are underpowered or conflicting, we should be transparent about the uncertainty rather than overstating the evidence. This might mean acknowledging that some widely performed procedures lack robust support, which could be uncomfortable but is ultimately necessary for the integrity of our profession and for good care. Dr. Leopold:Let’s pivot from the big picture to the smaller ones. Though I wish it were otherwise, I don’t think clinician-scientists are going to change their approach from frequentist statistics (p values) to Bayesian approaches tomorrow, so we need to offer clinicians who read journals an alternative approach today. I have offered a “Bayes Through the Back Door” approach (Fig. 2) that accounts for what is known before reading a study—which one might place in the context of the risks to patients that might accompany a proposed treatment—when choosing and interpreting p values in orthopaedic research. How does this seem to you? What do you see as its strengths and shortcomings? Dr. Shimozono: Your approach elegantly captures how experienced clinicians actually think. When I see a study with a p value of 0.03 that claims arthroscopic debridement improves ankle osteoarthritis symptoms, my skepticism isn't stubbornness but integration of prior knowledge—we know cartilage doesn't regenerate, and similar claims in the knee have been thoroughly debunked. The Bayesian framework's strength is legitimizing this clinical intuition while providing structure for its application. The main shortcoming is subjectivity in establishing prior probabilities. Two surgeons evaluating identical evidence might reach different conclusions based on their training, experience, and practice environment. This variability isn't necessarily problematic—it reflects genuine uncertainty in our field—but it could inadvertently reinforce preexisting biases. Additionally, the approach assumes a level of statistical literacy that may be optimistic. Many clinicians struggle with basic statistical concepts like confidence intervals, making informal Bayesian reasoning challenging unless we start to provide better statistical education in our training programs. Despite these limitations, your approach offers a valuable framework for more thoughtful evidence evaluation. It acknowledges that interpreting research findings isn't simply checking whether p < 0.05, but rather integrating multiple information sources—biological plausibility, clinical experience, and statistical evidence—weighted by their quality and relevance to the specific clinical context. Dr. Leopold:What would you offer as an alternative for the reader who recognizes the problem with the way clinical trialists (and journal editors) are handling this issue, to help that reader get more from what he or she reads in orthopaedic journals? Dr. Shimozono: I recommend a practical, checklist-driven approach. First, examine control group outcomes carefully. If control groups show meaningful improvement, the intervention must demonstrate substantially superior results to justify its risks and costs. Second, focus on clinical importance, not just statistical significance. Many “positive” findings in our analysis were only 2- to 5-point improvements on 100-point scales—differences too small for patients to perceive. Studies should always compare effect sizes to established MCIDs and calculate the number needed to treat. Third, ensure follow-up duration is adequate for the condition being outcomes can be if they improvements rather than In our analysis [7], over half of studies with findings would statistical significance the stricter 0.005 threshold, the of many Finally, consider When studies claim that established biological evidence should be Dr. one for our foot and ankle at the studies that would have been can you say what treatments what kinds of as a result of the In other words, which treatments for patients with foot and ankle treatments to in the RCTs you at the 0.05 level but would no be at the 0.005 Dr. Shimozono: the 52% of RCTs that would be according to our analysis [7], arthroscopic interventions for conditions were study ankle for osteoarthritis with p values between and would significance. This what with knee for as as for for ankle and would be reclassified as only evidence. p values between and about p < 0.05 and would significance entirely, which is given the irreversibility of these for and hallux valgus into this studies about significance. Studies of ankle protocols, and often p < our evidence is for conditions with outcomes than for conditions with The our study [7] is for conditions with natural on statistical than This doesn't mean these treatments but been recommending them based on evidence. Our findings should more and about uncertainty when it