科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Alpha psychiatry2026-08-01

A Cautious Integration With AI in the Clinic: A Standardized-Patient Pilot Study of ChatGPT's Reliability in Hamilton Depression Rating Scale Scoring.

Chun-Hung Chang, Szu-Wei Cheng, Wei-Jen Chen, Chung-Wen Chang, Ting-Hui Liu, Jia-Hau Lee, Sheng-Che Lin, Kuan-Pin Su

一句话结论 · In one sentence

ChatGPT demonstrated excellent agreement on total HAMD-21 scores in structured, text-based depression assessments, supporting the potential role of LLMs as adjunctive tools for standardized depression severity evaluation. However, item-level discrepancies and systematic scoring errors indicate that human oversight remains essential for clinically nuanced interpretation.

原始摘要(英文原文)· Original abstract
BACKGROUND: Artificial intelligence (AI) integration offers significant potential to improve mental healthcare, however, the reliability of large language models (LLMs) in performing nuanced clinical tasks remains an important and largely unanswered question. This study aimed to evaluate ChatGPT's performance in scoring the Hamilton Depression Rating Scale (HAMD-21) compared with expert raters using standardized patients (SPs). METHODS: Three senior mental health experts created and portrayed scenarios for ten SPs representing diverse depressive symptom profiles. Recorded interviews were transcribed and used as input for ChatGPT-4o. HAMD-21 scores generated by ChatGPT were compared with those assigned by expert raters and with predefined script-based reference scores. Inter-rater reliability was assessed using intraclass correlation coefficient (ICC), and differences between raters were evaluated using Steiger's tests. RESULTS: ChatGPT and the expert raters achieved good-to-excellent reliability for total HAMD-21 scores (experts: ICC = 0.9921; ChatGPT: ICC = 0.9739). However, expert raters achieved perfect ICCs on 11 individual items, whereas ChatGPT achieved perfect agreement on only 2 items. Steiger's test demonstrated that experts significantly outperformed ChatGPT on 10 individual items as well as on total scores (Z = 1.931, p = 0.0268). Qualitative review revealed that ChatGPT tended to overestimate scores on items related to insomnia and somatic symptoms (items 4-6 and 13) and frequently miscalculated total scores. CONCLUSIONS: ChatGPT demonstrated excellent agreement on total HAMD-21 scores in structured, text-based depression assessments, supporting the potential role of LLMs as adjunctive tools for standardized depression severity evaluation. However, item-level discrepancies and systematic scoring errors indicate that human oversight remains essential for clinically nuanced interpretation.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

A Cautious Integration With AI in the Clinic: A Standardized-Patient Pilot Study of ChatGPT's Reliability in Hamilton Depression Rating Scale Scoring. — 科研速览 Science Skim