Aquiles Darghan, Ariel Iván Ruiz-Parra, Jorge Eduardo Caminos, Nair González, Kevin Darghan, Sofia Alexandra Caminos Cepada
The Bayesian Youden Index positions itself not merely as a measure of statistical fit but as a robust calibration tool that decomposes overall classifier performance into predictive confidence under realistic prevalence conditions. The findings caution against prematurely discarding results from imbalanced datasets or automatically applying balancing methods, advocating instead for a nuanced analysis in which imbalance, mediated by classifier characteristics, can sometimes enhance rather than hinder evaluated performance.
BACKGROUND: The evaluation of binary classifiers under real-world conditions requires metrics that account for the dependence of predictive performance on disease prevalence. Classical accuracy metrics, including the Youden index, condition on the true class and remain invariant to changes in class distribution, making them unable to reflect the practical question of whether a positive classification is trustworthy in a given deployment context.
METHODS: The Bayesian Youden Index is derived analytically as a prevalence-dependent extension of the classical Youden index, grounded in the equivalence between Bayesian and classical sensitivity and specificity. Its behavior is characterized across the full prevalence range, and uncertainty in all point estimates is quantified through parametric bootstrap with B = 2,000 replications. The framework is illustrated using HOMA-IR for insulin resistance detection (n = 93) and is applicable to any binary classification task with a known confusion matrix and operational prevalence.
RESULTS: The Bayesian Youden Index reaches a maximum of 0.913 (95% CI:0.823,0.997) at a prevalence of 0.334 (95% CI:0.027,0.575), compared to the constant classical value of 0.890 (95% CI:0.782,0.978). At the observed prevalence of 0.484 the Bayesian value is 0.898. Two intersection points were identified at which the Bayesian and classical formulations coincide: P 1 = 0.190 (95% CI:0.118,0.542) and P 2 = 0518 (95% CI:0.504,0.652), defining three distinct prevalence regimes with different implications for classifier deployment. The index exhibits asymmetric behavior across the prevalence range, showing that a classifier's inherent sensitivity-specificity profile can be strategically leveraged through prevalence-aware threshold selection to manage class imbalance without resorting to artificial balancing techniques.
CONCLUSIONS: The Bayesian Youden Index positions itself not merely as a measure of statistical fit but as a robust calibration tool that decomposes overall classifier performance into predictive confidence under realistic prevalence conditions. The findings caution against prematurely discarding results from imbalanced datasets or automatically applying balancing methods, advocating instead for a nuanced analysis in which imbalance, mediated by classifier characteristics, can sometimes enhance rather than hinder evaluated performance.