Burhan Başkan, Yusuf Evcimen, Özgür Bülent Timuçin, İffet Yarımağa, Celal Sağmış
This study aimed to comprehensively evaluate the performance of seven leading large language models (LLMs) from the 2024-2025 period on Turkish Medical Specialty Examination (TUS) ophthalmology questions, analyzing factors such as question type, chronology, and clinical area, while critically assessing the risk of data contamination and the technology's realistic potential in medical education. A set of 210 TUS ophthalmology questions (2015-2024) was presented to seven LLMs (OpenAI o1-preview, Claude 4.5 Sonnet, GPT-4o, Gemini 2.5 Pro, DeepSeek-R1, Llama 3 70B, Command R+) using a standardized protocol. Performance was statistically compared based on overall accuracy, question type (case-based vs. knowledge-based), chronological period (old: 2015-2019 vs. new: 2020-2024), and clinical subspecialty. All responses underwent blinded, expert qualitative coding for error type and confidence. OpenAI o1-preview achieved the highest overall accuracy (95.2%). Claude 4.5 Sonnet and GPT-4o demonstrated 100% accuracy on case-based questions. A significant performance gap existed between closed-source (average 91.0%) and open-source models (average 66.1%) (p