科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Muş Alparslan Üniversitesi Sağlık Bilimleri Dergisi2026-07-31· Turkish

Evaluating Large Language Models as Educational Tools: A Multi-Model Analysis of Performance on the Turkish Medical Specialty Examination (TUS) Ophthalmology Questions Özet:

Burhan Başkan, Yusuf Evcimen, Özgür Bülent Timuçin, İffet Yarımağa, Celal Sağmış

原始摘要(英文原文)· Original abstract
This study aimed to comprehensively evaluate the performance of seven leading large language models (LLMs) from the 2024-2025 period on Turkish Medical Specialty Examination (TUS) ophthalmology questions, analyzing factors such as question type, chronology, and clinical area, while critically assessing the risk of data contamination and the technology's realistic potential in medical education. A set of 210 TUS ophthalmology questions (2015-2024) was presented to seven LLMs (OpenAI o1-preview, Claude 4.5 Sonnet, GPT-4o, Gemini 2.5 Pro, DeepSeek-R1, Llama 3 70B, Command R+) using a standardized protocol. Performance was statistically compared based on overall accuracy, question type (case-based vs. knowledge-based), chronological period (old: 2015-2019 vs. new: 2020-2024), and clinical subspecialty. All responses underwent blinded, expert qualitative coding for error type and confidence. OpenAI o1-preview achieved the highest overall accuracy (95.2%). Claude 4.5 Sonnet and GPT-4o demonstrated 100% accuracy on case-based questions. A significant performance gap existed between closed-source (average 91.0%) and open-source models (average 66.1%) (p
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Evaluating Large Language Models as Educational Tools: A Multi-Model Analysis of Performance on the Turkish Medical Specialty Examination (TUS) Ophthalmology Questions Özet: — 科研速览 Science Skim