科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ medRxiv : the preprint server for health sciences2026-09-17· oncology

Validating LLM judges for automated oversight of patient communication.

Zidu Xu, Johnathan Zeng, Shuang Zhou, Zhihong Zhang, Thibault Heintz, Marion Tonneau, Arsalan Yaghoubi, Bingyang Ye, Vikram Goddla, Lisa Lehmann, Yu-Hui Chen, Elad Sharon, David E Kozono, Anna Revette, Julia Maues, Thelma Brown, Paul Catalano, Raymond H Mak, Dimitry Dligach, Danielle S Bitterman

原始摘要(英文原文)· Original abstract
LLMs are increasingly used to mediate patient communication, yet scalable evaluation of their safety, accuracy, and communication quality remains an open problem. LLM judges have emerged as automated evaluators, but whether they can holistically replicate human expert judgment is unvalidated. Informed consent for clinical trials presents a demanding case for such validation because it requires conveying complex information to lay audiences under ethical and safety constraints. We developed a stakeholder-informed seven-criterion evaluation rubric spanning safety, reliability, and communication quality. Clinician reference ratings showed strong interrater reliability across all criteria. We validated the rubric on the Informed CONsent Benchmark (ICON-Bench) and benchmarked 19 LLM judges across multiple implementation strategies. LLM judges achieved strong clinician agreement for safety screening and factual verification (Spearman ρ > 0.80) but weaker agreement for communication quality ( ρ < 0.60). Safety-specialized guard models underperformed general-purpose models. Patient advocates rated communication quality lower than both clinicians and LLM judges. These findings support LLM judges for scalable patient communication oversight while demonstrating the need for recalibration to patient-centered evaluation standards.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Validating LLM judges for automated oversight of patient communication. — 科研速览 Science Skim