科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Surgery today2026-08-25

A Rasch-based analysis comparing the performance of updated multimodal large language models and surgeons on the Japanese surgical specialist examination.

Yuji Miyamoto, Takeshi Nakaura, Ayane Kawata, Toshinori Hirai, Hidetoshi Eguchi, Akinobu Taketomi, Masaaki Iwatsuki

一句话结论 · In one sentence

For this previously analyzed item set, updated multimodal LLMs achieved an overall accuracy below the surgeon benchmark, with the largest deficits on image-based items, supporting supervised and domain-specific use.

原始摘要(英文原文)· Original abstract
PURPOSE: To compare four multimodal large language models (LLMs) with surgeons on the 2023 Japanese Surgical Specialist Examination using item-level surgeon correct answer rates as the benchmark. METHODS: In this retrospective cross-sectional study, GPT-4.1, Claude Opus 4, Gemini 2.5 Pro, and o3 Pro were evaluated using 98 valid multiple-choice items, including 43 image-based and 55 text-only questions. The accuracy was examined overall, by image presence, and by subspecialty. Physician-anchored Rasch modeling placed LLMs and surgeons on a common latent scale, and logistic regression assessed the association between surgeon accuracy and LLM correctness at the item-level. RESULTS: The accuracy ranged from 77.6% for GPT-4.1 to 85.7% for o3 Pro. All models showed lower accuracy on image-based items than on text-only items. A Rasch analysis showed that all LLMs remained below the surgeons' overall, with relative abilities ranging from - 1.29 to - 0.74. Performance varied according to difficulty and subspecialty. Gastroenterology was consistently the weakest domain, whereas some models matched or exceeded the surgeon benchmark in selected areas, including respiratory, pediatrics, breast/endocrine, and emergency/anesthesiology. CONCLUSIONS: For this previously analyzed item set, updated multimodal LLMs achieved an overall accuracy below the surgeon benchmark, with the largest deficits on image-based items, supporting supervised and domain-specific use.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

A Rasch-based analysis comparing the performance of updated multimodal large language models and surgeons on the Japanese surgical specialist examination. — 科研速览 Science Skim