科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ JMIR AI2026-09-23

Implicit Bias in Large Language Model Diagnosis of Eating Disorders: Experimental Vignette Study.

Deija McCalla, Bochen Li, Saul Jaeger, Justin Jacques, Leo Gonzalez, Charles Silber, Cass Dykeman

一句话结论 · In one sentence

These findings demonstrate that LLMs exhibit systematic demographic biases in psychiatric diagnosis even when clinical content is held constant, revealing measurable patterns that can inform improvements to training data, model architecture, and clinical deployment frameworks.

原始摘要(英文原文)· Original abstract
BACKGROUND: Large language models (LLMs) are increasingly deployed in mental health applications, yet growing evidence suggests they encode algorithmic biases that influence clinical outputs. Because these models now mediate patient-facing decisions, such biases carry the potential for direct harm. Whether they systematically affect psychiatric diagnosis across demographic groups remains underexplored. OBJECTIVE: This study aims to examine whether LLMs exhibit implicit demographic biases when generating psychiatric diagnoses. METHODS: We developed 1152 synthetic clinical vignettes using a matched-pair design that manipulated gender, race and ethnicity, age, socioeconomic status, English proficiency, and urbanicity while holding clinical content constant. Vignettes were divided into control (unambiguous anorexia nervosa [AN]) and ambiguous conditions designed to permit differential diagnosis. Ten LLM configurations across 5 model families were tested. RESULTS: Control vignettes produced near-unanimous AN diagnoses (mean 100%, SD 0.1%), while ambiguous vignettes elicited greater variability (mean 23.6%, SD 10.1%). Intermodel agreement was moderate for ambiguous vignettes (Fleiss κ=0.410, 95% CI 0.397-0.422). Mixed-effects logistic regression with LLM as a random intercept revealed significant demographic biases: Black patients were over 6 times more likely to receive a major depressive disorder (MDD) diagnosis than White patients with identical presentations (odds ratio [OR] 6.09, 95% CI 5.13-7.24), Latine patients were over 9 times more likely (OR 9.57, 95% CI 8.00-11.45), and Asian patients were nearly 3 times more likely to receive an AN diagnosis (OR 2.88, 95% CI 2.44-3.42). Female patients were less likely than males to be diagnosed with AN (OR 0.43, 95% CI 0.37-0.49). CONCLUSIONS: These findings demonstrate that LLMs exhibit systematic demographic biases in psychiatric diagnosis even when clinical content is held constant, revealing measurable patterns that can inform improvements to training data, model architecture, and clinical deployment frameworks.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Implicit Bias in Large Language Model Diagnosis of Eating Disorders: Experimental Vignette Study. — 科研速览 Science Skim