科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ npj Health Systems2025-11-03· Standardization

Human evaluation of large language models in healthcare: gaps, challenges, and the need for standardization

Raghav Awasthi, Atharva Bhattad, S. Ramachandran, Shreya Mishra, Ashish K. Khanna, Jacek B. Cywiński, Kamal Maheshwari, Dwarikanath Mahapatra, Izabella DiRosa, Anabelle Cohen, Hajra Arshad, Aarit Atreja, Asma Alshukaili, Aryan Vohra, Nishant Singh, Francis Papay, Ashish Atreja, Rahul Kashyap, Piyush Mathur

原始摘要(英文原文)· Original abstract
Publications related to experimentation with Large Language Models (LLMs) in healthcare are rapidly increasing. While human evaluation remains the gold standard for evaluating LLMs, there is still a lack of standardization in its implementation. In this review article, we systematically examine studies involving LLMs in healthcare that have conducted human evaluations. We analyze the metrics used, assess their variability across studies. We also propose a standardized framework along with an interactive open web application HumanELY, to facilitate human evaluation. We believe that use of HumanELY will provide an opportunity for consistent, comprehensive, reliable, reproducible, and measurable human evaluations of LLM in healthcare. HumanELY is publicly available at https://www.brainxai.com/humanely .
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Human evaluation of large language models in healthcare: gaps, challenges, and the need for standardization — 科研速览 Science Skim