科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ npj Digital Medicine2025-10-07· Transparency (behavior)

The evaluation illusion of large language models in medicine

Monica Agrawal, Irene Y. Chen, Freya Gulamali, Shalmali Joshi

原始摘要(英文原文)· Original abstract
While large language models (LLMs) hold promise for transforming clinical healthcare, current comparisons and benchmark evaluations of large language models in medicine often fail to capture real-world efficacy. Specifically, we highlight how key discrepancies arising from choices of data, tasks, and metrics can limit meaningful assessment of translational impact and cause misleading conclusions. Therefore, we advocate for rigorous, context-aware evaluations and experimental transparency across both research and deployment.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

The evaluation illusion of large language models in medicine — 科研速览 Science Skim