科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ JMIR formative research2026-09-02

Performance of ChatGPT, Claude, and AMBOSS on the European Board of Urology In-Service Assessment and Alignment With the European Association of Urology 2025 Guidelines: Comparative Study.

Mohamed Eldaneen, Shaza Hendy, Omar Ramadan, Panagiotis Nikolinakos, Asif Muneer, Hussain M Alnajjar, Karl H Pang

一句话结论 · In one sentence

AI tools can achieve high accuracy on EBU-style assessments; however, differences in reasoning quality and guideline adherence are evident. ChatGPT demonstrated superior performance across all evaluated domains, supporting its role as a potential adjunct in postgraduate urology education.

原始摘要(英文原文)· Original abstract
BACKGROUND: Recent advances in AI, particularly large language models, have generated growing interest in their application to medical education and examination preparation. However, the accuracy, reasoning quality, and adherence to clinical guidelines of these tools in postgraduate urology assessments remain unclear. OBJECTIVE: This study aimed to evaluate the performance of 3 AI tools, ChatGPT (GPT-4.0), Claude (version 4.5), and AMBOSS, on European Board of Urology (EBU)-style multiple-choice questions, with a particular focus on accuracy, insight, concordance, and adherence to European Association of Urology (EAU) guidelines. METHODS: A total of 200 single-best-answer questions from the EBU In-Service Assessment workbook (2021-2022) were input into each AI model. Models were prompted to select an answer and provide an explanation. Two urologists with post-Fellowship of the Royal College of Surgeons (FRCS) training independently assessed the outputs. Accuracy was defined as correct answer selection. Concordance was defined as the logical alignment between the answer and its explanation. Insight was evaluated across 3 domains-nonobvious deduction, discriminative reasoning, and clinical validity-and was graded as low, moderate, or high. RESULTS: ChatGPT demonstrated the highest accuracy (171/200, 85.5%), compared to Claude and AMBOSS (both 159/200, 79.5%; P=.14). Concordance was also significantly higher for ChatGPT (190/200, 95%) than for Claude (176/200, 88%) and AMBOSS (152/200, 76%; P<.001). Nonobvious deduction was predominantly low to moderate across all models, reflecting the recall-based nature of many questions. ChatGPT and Claude showed stronger discriminative reasoning, while AMBOSS demonstrated limited exclusion of alternative options. Clinical validity was high overall, with ChatGPT showing the greatest consistency with EAU guidelines. There was substantial agreement between the 2 reviewers (weighted κ coefficient >0.61). CONCLUSIONS: AI tools can achieve high accuracy on EBU-style assessments; however, differences in reasoning quality and guideline adherence are evident. ChatGPT demonstrated superior performance across all evaluated domains, supporting its role as a potential adjunct in postgraduate urology education.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Performance of ChatGPT, Claude, and AMBOSS on the European Board of Urology In-Service Assessment and Alignment With the European Association of Urology 2025 Guidelines: Comparative Study. — 科研速览 Science Skim