Cedric Neumann, Andrew Anthony Matas, Jeffrey Jamison Geppert
The selection and use of quality measures in value-based payment rely on single-period methods that conflate sampling noise with genuine within-provider year-to-year execution variability, producing assessments that are unstable across years and dependent on patient volume. We introduce a Longitudinal Beta-Binomial model, which separates three sources of variability simultaneously and produces a year-invariant measure-level assessment τ and a volume-independent provider-specific assessment τ n . Applied to five years of data (program years 2021-2025) for 744 providers on two CMS mental health quality measures, the analysis yields a striking reversal: under the standard Beta-Binomial framework, FUH-30 appears far more discriminating than READM-30-IPF ( ρ ( j ) = 0 . 78 -0.87 vs 0.51-0.63 per year); separating execution variability from sampling noise reverses this conclusion ( τ = 0 . 621 [0.598, 0.642] for FUH-30 vs τ = 0 . 758 [0.730, 0.784] for READM-30-IPF, 95% posterior credible intervals), because FUH-30's year-to-year provider instability is three times larger than its sampling noise. The framework proposes three requirements simultaneously: discriminability ( τ can separate good from poor health-care providers and is year-invariant where ρ ( j ) swings substantially across years); fairness (both τ and τ n are independent of patient volume, removing the structural disadvantage faced by small and rural providers); and actionability ( τ n distinguishes consistently substandard from erratic providers, enabling targeted intervention and more honest measure adoption decisions).