Wei Xiao, Huanyu Zhang, Xiaoyi Chen, Jun Cai, Xuchen Luo, Jiehua Deng
The three models answered many late-life depression questions accurately and safely, but none demonstrated consistently reliable performance across complex or high-risk scenarios. Clinically important weaknesses involved geriatric-specific interpretation, recognition of indirect risk signals, crisis-response completeness, under-triage, and response stability. General-purpose large language models may support selected low-risk educational tasks but should not independently guide emergency triage, medication changes, suicide-risk management, or other safety-critical decisions in older adults.
BACKGROUND: General-purpose large language models are increasingly used by patients and caregivers to obtain mental health information and guidance about when professional care is required. In late-life depression, broadly accurate information may nevertheless be unsafe when cognitive change, multimorbidity, frailty, polypharmacy, self-neglect, caregiver dependence, or suicide risk is not adequately recognized. This study compared the clinical accuracy, safety, geriatric-specific appropriateness, triage performance, and response consistency of three large language models when answering patient- and caregiver-centered questions about late-life depression.
METHODS: We conducted a blinded, paired benchmarking study using 90 questions covering six geriatric psychiatry domains and equally distributed across low-, moderate-, and high-risk strata. Each question was submitted independently to GPT-5.5 Instant via ChatGPT, Gemini 3.5 Flash via Gemini, and Seed2.0 Pro via Doubao, generating 270 primary-round responses. A stratified subset of 30 questions was resubmitted in separate conversations to assess test-retest consistency, yielding 360 responses overall. Two psychiatrists independently evaluated anonymized outputs against prespecified item-specific reference standards, with clinically important disagreements adjudicated by a third senior psychiatrist. The primary outcome was the proportion of clinically acceptable responses, defined using accuracy, clinical safety, geriatric appropriateness, warning-sign recognition, and triage criteria.
RESULTS: Clinically acceptable responses were generated for 78.9% of questions by ChatGPT, 72.2% by Gemini, and 60.0% by Doubao (overall P<0.001). The difference between ChatGPT and Gemini was not statistically significant, whereas both outperformed Doubao. This model ranking remained unchanged under alternative core-safety and more stringent optimal-response definitions. Complete geriatric-specific appropriateness was achieved in 64.4%, 55.6%, and 43.3% of responses, respectively. Major safety errors occurred in 5.6% of ChatGPT responses, 10.0% of Gemini responses, and 16.7% of Doubao responses (raw P = 0.015; FDR-adjusted q=0.023). Performance declined substantially with increasing clinical risk. In exploratory post hoc caregiver-centered analyses, clinical acceptability was 76.7% for ChatGPT, 70.0% for Gemini, and 53.3% for Doubao, while explicit caregiver-directed action was present in 80.0%, 70.0%, and 53.3% of responses, respectively. Among high-risk questions, clinically acceptable response rates were 63.3%, 50.0%, and 36.7%, while under-triage occurred in 16.7%, 26.7%, and 36.7% of responses, respectively. Test-retest clinical consistency was highest for ChatGPT (90.0%), followed by Gemini (83.3%) and Doubao (73.3%).
CONCLUSIONS: The three models answered many late-life depression questions accurately and safely, but none demonstrated consistently reliable performance across complex or high-risk scenarios. Clinically important weaknesses involved geriatric-specific interpretation, recognition of indirect risk signals, crisis-response completeness, under-triage, and response stability. General-purpose large language models may support selected low-risk educational tasks but should not independently guide emergency triage, medication changes, suicide-risk management, or other safety-critical decisions in older adults.