Kwang Hwan ParK, Dong Woo Shim, Jin Woo Lee, Hak Jun Kim, Dong Hun Suh, Yeong Jeon, Ji Hun Park, Gi Won Choi
LLM performance for foot and ankle disorders differed substantially across models. GPT-5 Thinking and GPT-o3 exhibited superior performance under the tested conditions, emphasizing the necessity of targeted model selection for clinical decision-making. However, because LLMs evolve rapidly, these findings should be interpreted as a time-specific snapshot rather than as fixed or generalizable rankings of model performance.
BACKGROUND: This study compared seven large language models (LLMs) to identify the optimal model for foot and ankle clinical decision support.
METHODS: The LLMs answered 20 multiple-choice questions (MCQs) and 20 open-ended clinical questions. MCQ accuracy was recorded. Three blinded foot and ankle surgeons evaluated the open-ended responses based on accuracy, completeness, and clinical relevance (total score range: 3-21).
RESULTS: GPT-o3 and GPT-5 Thinking achieved the highest MCQ accuracy (95%). For open-ended evaluations, mean total scores differed significantly across the models (p < 0.001). GPT-5 Thinking (19.98 ± 0.66) and GPT-o3 (19.67 ± 0.51) attained the highest scores, significantly outperforming the other five models, with no statistical difference between these top two performers.
CONCLUSION: LLM performance for foot and ankle disorders differed substantially across models. GPT-5 Thinking and GPT-o3 exhibited superior performance under the tested conditions, emphasizing the necessity of targeted model selection for clinical decision-making. However, because LLMs evolve rapidly, these findings should be interpreted as a time-specific snapshot rather than as fixed or generalizable rankings of model performance.