Yuji Miyamoto, Takeshi Nakaura, Ayane Kawata, Toshinori Hirai, Hidetoshi Eguchi, Akinobu Taketomi, Masaaki Iwatsuki
For this previously analyzed item set, updated multimodal LLMs achieved an overall accuracy below the surgeon benchmark, with the largest deficits on image-based items, supporting supervised and domain-specific use.
PURPOSE: To compare four multimodal large language models (LLMs) with surgeons on the 2023 Japanese Surgical Specialist Examination using item-level surgeon correct answer rates as the benchmark.
METHODS: In this retrospective cross-sectional study, GPT-4.1, Claude Opus 4, Gemini 2.5 Pro, and o3 Pro were evaluated using 98 valid multiple-choice items, including 43 image-based and 55 text-only questions. The accuracy was examined overall, by image presence, and by subspecialty. Physician-anchored Rasch modeling placed LLMs and surgeons on a common latent scale, and logistic regression assessed the association between surgeon accuracy and LLM correctness at the item-level.
RESULTS: The accuracy ranged from 77.6% for GPT-4.1 to 85.7% for o3 Pro. All models showed lower accuracy on image-based items than on text-only items. A Rasch analysis showed that all LLMs remained below the surgeons' overall, with relative abilities ranging from - 1.29 to - 0.74. Performance varied according to difficulty and subspecialty. Gastroenterology was consistently the weakest domain, whereas some models matched or exceeded the surgeon benchmark in selected areas, including respiratory, pediatrics, breast/endocrine, and emergency/anesthesiology.
CONCLUSIONS: For this previously analyzed item set, updated multimodal LLMs achieved an overall accuracy below the surgeon benchmark, with the largest deficits on image-based items, supporting supervised and domain-specific use.