科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Foot and ankle surgery : official journal of the European Society of Foot and Ankle Surgeons2026-08-11

Comparative performance of large language models for foot and ankle clinical decision support: A multi-rater evaluation of seven models.

Kwang Hwan ParK, Dong Woo Shim, Jin Woo Lee, Hak Jun Kim, Dong Hun Suh, Yeong Jeon, Ji Hun Park, Gi Won Choi

一句话结论 · In one sentence

LLM performance for foot and ankle disorders differed substantially across models. GPT-5 Thinking and GPT-o3 exhibited superior performance under the tested conditions, emphasizing the necessity of targeted model selection for clinical decision-making. However, because LLMs evolve rapidly, these findings should be interpreted as a time-specific snapshot rather than as fixed or generalizable rankings of model performance.

原始摘要(英文原文)· Original abstract
BACKGROUND: This study compared seven large language models (LLMs) to identify the optimal model for foot and ankle clinical decision support. METHODS: The LLMs answered 20 multiple-choice questions (MCQs) and 20 open-ended clinical questions. MCQ accuracy was recorded. Three blinded foot and ankle surgeons evaluated the open-ended responses based on accuracy, completeness, and clinical relevance (total score range: 3-21). RESULTS: GPT-o3 and GPT-5 Thinking achieved the highest MCQ accuracy (95%). For open-ended evaluations, mean total scores differed significantly across the models (p < 0.001). GPT-5 Thinking (19.98 ± 0.66) and GPT-o3 (19.67 ± 0.51) attained the highest scores, significantly outperforming the other five models, with no statistical difference between these top two performers. CONCLUSION: LLM performance for foot and ankle disorders differed substantially across models. GPT-5 Thinking and GPT-o3 exhibited superior performance under the tested conditions, emphasizing the necessity of targeted model selection for clinical decision-making. However, because LLMs evolve rapidly, these findings should be interpreted as a time-specific snapshot rather than as fixed or generalizable rankings of model performance.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Comparative performance of large language models for foot and ankle clinical decision support: A multi-rater evaluation of seven models. — 科研速览 Science Skim