Chun-Hung Chang, Szu-Wei Cheng, Wei-Jen Chen, Chung-Wen Chang, Ting-Hui Liu, Jia-Hau Lee, Sheng-Che Lin, Kuan-Pin Su
ChatGPT demonstrated excellent agreement on total HAMD-21 scores in structured, text-based depression assessments, supporting the potential role of LLMs as adjunctive tools for standardized depression severity evaluation. However, item-level discrepancies and systematic scoring errors indicate that human oversight remains essential for clinically nuanced interpretation.
BACKGROUND: Artificial intelligence (AI) integration offers significant potential to improve mental healthcare, however, the reliability of large language models (LLMs) in performing nuanced clinical tasks remains an important and largely unanswered question. This study aimed to evaluate ChatGPT's performance in scoring the Hamilton Depression Rating Scale (HAMD-21) compared with expert raters using standardized patients (SPs).
METHODS: Three senior mental health experts created and portrayed scenarios for ten SPs representing diverse depressive symptom profiles. Recorded interviews were transcribed and used as input for ChatGPT-4o. HAMD-21 scores generated by ChatGPT were compared with those assigned by expert raters and with predefined script-based reference scores. Inter-rater reliability was assessed using intraclass correlation coefficient (ICC), and differences between raters were evaluated using Steiger's tests.
RESULTS: ChatGPT and the expert raters achieved good-to-excellent reliability for total HAMD-21 scores (experts: ICC = 0.9921; ChatGPT: ICC = 0.9739). However, expert raters achieved perfect ICCs on 11 individual items, whereas ChatGPT achieved perfect agreement on only 2 items. Steiger's test demonstrated that experts significantly outperformed ChatGPT on 10 individual items as well as on total scores (Z = 1.931, p = 0.0268). Qualitative review revealed that ChatGPT tended to overestimate scores on items related to insomnia and somatic symptoms (items 4-6 and 13) and frequently miscalculated total scores.
CONCLUSIONS: ChatGPT demonstrated excellent agreement on total HAMD-21 scores in structured, text-based depression assessments, supporting the potential role of LLMs as adjunctive tools for standardized depression severity evaluation. However, item-level discrepancies and systematic scoring errors indicate that human oversight remains essential for clinically nuanced interpretation.