科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Science2026-04-30· Computer science

Performance of a large language model on the reasoning tasks of a physician

Peter G. Brodeur, Thomas A Buckley, Zahir Kanjee, Ethan Goh, Evelyn Ling, Priyank Jain, Stephanie Cabral, Raja-Elie Abdulnour, Adrian D. Haimovich, Jason A. Freed, Andrew Olson, Daniel J Morgan, Jason Hom, Robert Gallo, Liam G McCoy, Haadi Mombini, Christopher Lucas, M. Fotoohi, Matthew Gwiazdon, Daniele Restifo, Daniel Restrepo, Eric Horvitz, Jonathan Chen, Arjun K. Manrai, Adam Rodman

原始摘要(英文原文)· Original abstract
More than 65 years ago, complex clinical diagnostic reasoning cases were introduced as the gold standard for the evaluation of expert medical computing systems, a standard that has held ever since. In this study, we report the results of a physician evaluation of a large language model (LLM) on challenging clinical cases across five experiments with a baseline of hundreds of physicians. We then report a real-world study comparing human expert and artificial intelligence (AI) second opinions in randomly selected patients in the emergency room of a major tertiary academic medical center. In all experiments, the LLM outperformed physician baselines and displayed continued improvement from prior generations of AI clinical decision support. Our study suggests that LLMs have eclipsed most benchmarks of clinical reasoning, motivating the urgent need for prospective trials.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Performance of a large language model on the reasoning tasks of a physician — 科研速览 Science Skim