Angélica Atehortúa, Uma-Maria Lal-Trehan Estrada, Paula Petrone
BackgroundDespite established sex differences in disease risk and clinical presentation, most clinical machine-learning models treat sex as a simple covariate without explicitly incorporating detailed sex-specific clinical history. Whether such information improves predictive performance, clinical utility, and between-sex fairness in multimorbid populations remains unclear. MethodsUsing baseline data from the PLCO cohort (≈155,000 participants), we developed XGBoost models for prevalent disease identification by comparing a reference model based on shared predictors and biological sex with an augmented model incorporating sex-specific clinical history. Leave-one-center-out internal-external validation was performed across multimorbidity subgroups. Discrimination, calibration, clinical utility, and between-sex fairness were evaluated. A structural-missingness ablation was additionally performed to distinguish the contribution of the original sex-specific clinical values from that of their missingness pattern. ResultsSex-specific clinical history improved discrimination and calibration, with larger gains at higher multimorbidity burden. In the three-disease group, the augmented model achieved an AUC of 0.814 (95% CI 0.804-0.826), a PR-AUC of 0.624 (95% CI 0.577-0.672), and a Brier score of 0.172 (95% CI 0.165-0.180), outperforming the reference model. Decision-analytic utility was negligible in the single-disease setting but increased with multimorbidity burden. Structural-missingness ablation showed negligible gains from the missingness pattern alone (ΔAUC = 0.00017), whereas restoring the original sex-specific values yielded substantially greater improvement (ΔAUC = 0.01019). Performance gains did not consistently reduce between-sex disparities, with effects varying across disease phenotypes and multimorbidity settings. Conclusions Sex-specific clinical history provides clinically relevant predictive information beyond predictors shared across women and men, with the greatest value observed in more complex multimorbid settings. However, improvements in predictive performance and clinical utility do not necessarily translate into reduced between-sex disparities. These findings support joint evaluation of predictive performance, calibration, clinical utility, and fairness when assessing the value of sex-specific information in clinical AI.