科研速览继续刷下去 →
◆ The American journal of geriatric psychiatry : official journal of the American Association for Geriatric Psychiatry2026-08-22

Systemic Failure Modes of Large Language Models in Geriatric Psychiatric Assessment: A Scoping Review.

Yu-Shiou Lin, Yun-Ting Lee

一句话结论

Current LLM architectures remain insufficient for autonomous geriatric psychiatric assessment. Near-term use should be limited to low-risk support functions with human oversight, structured cognitive forcing safeguards, and evaluation protocols that stress-test longitudinal complexity, adversarial conditions, multimodal safety events, and bias.

原始摘要(原文)
BACKGROUND: Large language models (LLMs) show strong performance on medical examinations, yet their reliability in real-world psychogeriatric assessment is uncertain, where delirium, dementia, and late-life depression often coexist amid frailty, polypharmacy, and long clinical histories. We aimed to map empirically evaluated LLM failure modes most relevant to geriatric psychiatric assessment and risk management. METHODS: We conducted a systematic scoping review of empirical evaluations published between January 1, 2023, and December 31, 2025. We searched biomedical and technical sources and included studies that tested LLMs on clinically relevant tasks, including diagnostic reasoning, longitudinal history integration, safety and medication reasoning, and bias-related outcomes. Evidence was charted and synthesized qualitatively to derive a pragmatic taxonomy of recurrent failure patterns. RESULTS: Forty-seven studies met inclusion criteria. Four convergent failure modes emerged: (1) diagnostic instability in longitudinal reasoning, including degraded retrieval within long contexts and loss of mid-history cues; (2) adversarial vulnerability, with elevated hallucination under misleading inputs and flawed rationales despite correct final answers; (3) multimodal temporal sparsity in video-based models that can omit brief safety-critical events such as falls; and (4) systemic ageism and value misalignment that can distort clinical narratives, risk estimation, and care recommendations. CONCLUSIONS: Current LLM architectures remain insufficient for autonomous geriatric psychiatric assessment. Near-term use should be limited to low-risk support functions with human oversight, structured cognitive forcing safeguards, and evaluation protocols that stress-test longitudinal complexity, adversarial conditions, multimodal safety events, and bias.
读原文 ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文

Systemic Failure Modes of Large Language Models in Geriatric Psychiatric Assessment: A Scoping Review. — 科研速览 Science Skim