科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ npj Mental Health Research2026-07-09· Benchmark (surveying)

PsyEval: a comprehensive large language model evaluation benchmark for mental health

Haoan Jin, Siyuan Chen, Dilawaier Dilixiati, Yewei Jiang, Kenny Q. Zhu, Mengyue Wu

原始摘要(英文原文)· Original abstract
Evaluating large language models (LLMs) in the mental health domain presents distinct challenges due to the subtle, context-dependent, and subjective nature of psychological symptoms. We introduce PsyEval, a benchmark specifically designed to evaluate LLMs in mental health-related tasks across three core dimensions: knowledge, diagnosis, and emotional support. PsyEval is constructed to reflect the complexity of mental health scenarios and provides a structured framework for assessing model performance within this sensitive domain. Using PsyEval, we evaluate eleven advanced LLMs with different prompting strategies to investigate how prompting affects their responses. The results reveal considerable gaps in LLMs' current ability to reason accurately and respond appropriately in mental health contexts, while also indicating promising directions for future model enhancement.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

PsyEval: a comprehensive large language model evaluation benchmark for mental health — 科研速览 Science Skim