科研速览 · Science Skim继续刷下去 · Keep skimming →
◇ medRxiv2026-08-19· psychiatry and clinical psychology

Clinically Grounded AI-Scribing in Psychotherapy: Benchmarking LLMs Against Expert Documentation in the iCARE Framework

P. K. Adhikary, S. Singh, S. Singh, P. Sharma, P. Soni, R. Choudhary, C. Saxena, P. Chauhan, S. K. Gupta, K. S. Deb, S. M. Singh, T. Chakraborty

原始摘要(英文原文)· Original abstract
Background: AI-scribes based on large language models (LLMs) are being rapidly adopted in healthcare, yet their use in psychotherapy remains limited because most systems produce brief, administratively oriented notes that miss the interpretative and contextual nuance essential to mental-health documentation. We argue that AI documentation for psychotherapy should instead be built on a comprehensive structure of history, clinical evaluation, and risk assessment that can subsequently be condensed into any brief format. Objective: To this end we introduce iCARE (Identifying information, Chief concerns and clinical history, Assessment and analysis, Risk identification, Evaluation of progress and action plan), a structured, clinically grounded, 17-section documentation format developed through iterative review by psychiatrists and clinical psychologists. Methods: Using the publicly available HOPE dataset, we created iHOPE, a corpus of 174 re-transcribed, speaker-diarized therapy sessions with expert-written gold-standard notes in the iCARE format. We benchmarked eleven contemporary LLMs (nine open-source and two proprietary) in zero-shot and one-shot settings, evaluating outputs with lexical (BLEU, METEOR, ROUGE-L) and semantic (BERTScore, BLEURT, InfoLM) metrics and with TRACE (Trustworthiness, Relevance, Accuracy, Comprehensiveness, Expression), a five-domain human-evaluation framework we designed for mental-health documentation. Results: Closed-source models consistently outperformed open-source ones, with GPT-4o-mini achieving the best overall automatic-evaluation scores and Gemini Pro excelling at interpretative sections such as crisis-marker identification. All models struggled with temporal reasoning, such as past-session and next-session details, and lexical overlap with human notes was low despite strong semantic alignment. In blinded expert review, clinical preference did not always mirror automatic benchmarks, with Mistral, a smaller open model, emerging as a surprise favourite. Conclusions: The iCARE format, iHOPE dataset, and TRACE framework together provide a foundation for clinically valid, therapist-centred AI documentation that assists rather than replaces clinicians.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Clinically Grounded AI-Scribing in Psychotherapy: Benchmarking LLMs Against Expert Documentation in the iCARE Framework — 科研速览 Science Skim