Eisuke Iwamoto, Kentaro Murase
Re-expressing effect sizes as clinical probabilities may provide a more intuitive interpretation of population-level evidence. This framework enables the interpretation of improvement, worsening, and stability simultaneously from a single effect size using complementary probability metrics. Although not intended for individual prediction, it may support evidence interpretation when patient-specific information is still limited. In routine clinical practice, however, evidence should ultimately be applied in conjunction with patient-specific information obtained through history taking and clinical examination. Future prospective studies using individual patient data are needed to validate this framework and evaluate its clinical applicability.
INTRODUCTION: Effect sizes such as mean differences (MDs) and standardized mean differences (SMDs) are widely used to summarize treatment effects. However, these measures do not directly indicate the likelihood of clinically meaningful improvement and may be difficult to interpret in clinical practice. This study presents an exploratory methodological framework for re-expressing effect sizes as clinically meaningful probabilities and examines whether these probabilities can support the interpretation and communication of evidence.
METHODS: We applied the proposed framework to published effect sizes from representative acupuncture randomized controlled trials and meta-analyses. Assuming normal distributions, MDs and SMDs were re-expressed as probabilities of improvement, worsening, and stability using minimal clinically important difference (MCID) thresholds. These probabilities were further expressed using odds ratios (ORs), therapeutic leverage ratios (TLRs), and absolute probability differences (ΔP). Sensitivity analyses were performed to examine the influence of MCID thresholds, effect sizes, baseline probabilities, and distributional assumptions on probability-based interpretations.
RESULTS: Previously reported MDs and SMDs could be re-expressed as probabilities of improvement, worsening, and stability. ORs, TLRs, and ΔP provided different perspectives on the same probability changes. Sensitivity analyses showed that absolute probabilities varied across assumptions, whereas the overall three-category structure of improvement, worsening, and stability was preserved across all scenarios.
CONCLUSION: Re-expressing effect sizes as clinical probabilities may provide a more intuitive interpretation of population-level evidence. This framework enables the interpretation of improvement, worsening, and stability simultaneously from a single effect size using complementary probability metrics. Although not intended for individual prediction, it may support evidence interpretation when patient-specific information is still limited. In routine clinical practice, however, evidence should ultimately be applied in conjunction with patient-specific information obtained through history taking and clinical examination. Future prospective studies using individual patient data are needed to validate this framework and evaluate its clinical applicability.