科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Foot & ankle orthopaedics2026-07-01

High Accuracy, Questionable References: Large Language Models' Performance and Citation Reliability in Foot and Ankle Surgery Examinations.

Benjamin Nieves-Lopez, Andrea Fabregas, Gonzalo F Del Rio Montesinos, José A Rosario-González, Jean L Santos-Agrait, Edward T Haupt, Keith T Aziz

一句话结论 · In one sentence

ChatGPT‑5.4 showed the highest accuracy on foot and ankle SAE questions, with diminished image-related performance gaps relative to prior generations, but also the highest rate of reference fabrication. All models produced fabricated citations, underscoring the need to verify AI-generated references and to use LLMs as adjunct rather than primary educational tools in foot and ankle surgery.

原始摘要(英文原文)· Original abstract
BACKGROUND: Artificial intelligence large language models (LLMs) are increasingly used for medical education and decision support in orthopaedics, yet their performance on foot and ankle surgery board-style questions remains underexplored. Additionally, the capability of LLMs to provide reliable references and source material for education has not been evaluated. Therefore, this study aimed to compare the accuracy and citation reliability of ChatGPT‑5.4, Gemini‑3, and Copilot on foot and ankle surgery Self-Assessment Examination (SAE) questions. METHODS: A total of 192 multiple-choice questions (64 text-only, 128 image-based) from OrthoBullets Foot and Ankle SAE forms D and E were entered into ChatGPT‑5.4, Gemini‑3, and Copilot; video-based items were excluded because of Copilot limitations. Accuracy was compared overall and by question type and imaging modality. References volunteered or prompted were classified as real, partially fabricated, or completely fabricated and compared across models and by answer correctness. Statistical significance was set at a P value <.05. RESULTS: ChatGPT‑5.4 achieved the highest overall accuracy (89.6%) vs Gemini‑3 (76.6%) and Copilot (74.0%). ChatGPT‑5.4 and Gemini‑3 showed accuracy that did not differ significantly on text-only vs image-based questions, whereas Copilot performed significantly worse on image-based items (P = .028). With the numbers available, no significant difference in accuracy could be detected by imaging modality for any model. ChatGPT‑5.4, Gemini‑3, and Copilot fabricated 22.9%, 16.1%, and 2.8% of all references, respectively; most ChatGPT‑5.4 and Gemini‑3 fabrications were partially fabricated, whereas all Copilot fabrications were completely fabricated. ChatGPT‑5.4 fabricated significantly more references when it answered incorrectly (P = .013). CONCLUSION: ChatGPT‑5.4 showed the highest accuracy on foot and ankle SAE questions, with diminished image-related performance gaps relative to prior generations, but also the highest rate of reference fabrication. All models produced fabricated citations, underscoring the need to verify AI-generated references and to use LLMs as adjunct rather than primary educational tools in foot and ankle surgery.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

High Accuracy, Questionable References: Large Language Models' Performance and Citation Reliability in Foot and Ankle Surgery Examinations. — 科研速览 Science Skim