Chiara Barbati, Virginia Casigliani, Caterina Rizzo, Anna Odone
Patients are increasingly turning to large language models for health information, which makes the consistency of these systems' outputs across patient groups a public health concern. Demographic bias is harder to see than fabrication because it escapes accuracy benchmarks and shows up instead in the language the models produce. Gender bias is the form most consistently reported, with effects on clinical documentation and reasoning that often compound with ethnicity, socioeconomic status and other attributes. Drawing on evidence from gender medicine, we argue that it calls for sex- and gender-disaggregated evaluation, gender-medicine expertise in benchmark design and stratified monitoring after deployment. The methods needed for this already exist. What is still unsettled is whether gender equity will be treated as a design requirement or left until its consequences can no longer be ignored.