S. He, J. W. Joseph, P. Safari, N. Akomeh, A. Goff, P. J. Slusarz, A. Mohamed, S. Lord, J. N. Goldstein, A. S. Raja, C. Kabrhel, B. W. Locke, C. Rohlfsen, D. M. Liebovitz
Objectives To quantify complementary meanings of binary clinical calculator results: information beyond the observed outcome frequency, post-result risk, information direction, and precision. Design Catalogue-based cross-sectional analysis of published aggregate performance data, with a reproducible row-level audit. Setting The 847-calculator MDCalc catalogue and a fixed collection of publicly available reports assembled without a systematic literature search. Participants Evaluations comparing a binary calculator result with a binary clinical outcome using a complete or reconstructable 2 x 2 table. Main outcome measures Percentage reduction in uncertainty and corresponding information gain in bits; risk after positive and negative classifications; the proportion of information from each classification; and whether the upper 95% confidence bound for post-negative risk supported specified thresholds. Results The analysis included 482 evaluations of 407 calculators. The median study size was 422 and the median observed outcome frequency was 17.6%. Results reduced uncertainty by a median 14.9% (interquartile range 6.2%-30.5%; 95% confidence interval 12.1% to 17.2%), corresponding to 0.093 bits. In 329 evaluations (68.3%), one classification supplied more than 60% of average information. PERC reduced uncertainty by 4.3% on average while its negative classification lowered observed risk from 7.6% to 1.0%. Two HEART thresholds in the same cohort shifted the positive-classification information share from 39.1% to 79.8%. Post-negative risk was below 2% in 157 evaluations by point estimate, but the upper 95% confidence bound was below 2% in only 67; 90 of 157 (57%) did not support the apparent threshold. Conclusions Clinical calculator results have no single quantitative meaning. Conventional performance, information added beyond the observed outcome frequency, post-result risk, classification-specific information, and precision provide complementary interpretations. Whether acting on a result improves care remains a separate question.