科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ BMC Oral Health2025-12-03· Endodontics

Comparison of chatbots’ accuracy in endodontics questions in dentistry specialization exam in Türkiye: ChatGPT-4o, Gemini Advanced, Copilot, and Claude

Ceren Turan Gökduman, Esra Arılı Öztürk, Şule Aktaş, Burhan Can Çanakçı

原始摘要(英文原文)· Original abstract
BACKGROUND: The aim of this study was to evaluate the performance of various chatbots (ChatGPT-4o, Gemini Advanced, Microsoft Copilot, and Claude) on endodontic questions in the Turkish Dentistry Specialization Exam (DUS) based on topic, year, and Bloom’s taxonomy. Additionally, the study aimed to contribute to the use of artificial intelligence chatbots as a supplementary learning tool by evaluating their reliability and limitations in dental education. METHODS: The accuracy rates of ChatGPT-4o, Gemini Advanced, Microsoft Copilot, and Claude were assessed using 122 endodontic questions from the DUS between 2012 and 2021. Each question was given to the chatbots in a new chat. Accuracy rates were compared based on year, topic, and Bloom’s taxonomy. RESULTS: The accuracy rates of the chatbots were similar: 82.8% for ChatGPT-4o, 83.6% for Gemini Advanced, 77.9% for Copilot, and 82.8% for Claude. Additionally, the accuracy of Copilot and Claude was significantly higher during 2012–2015 compared to 2016–2021 (p = 0.020; p = 0.018). According to Bloom’s taxonomy, ChatGPT-4o, Copilot, and Claude showed higher accuracy in 2012–2015 than in 2016–2021 on low-level questions (p = 0.040; p = 0.011; p = 0.005). CONCLUSIONS: Within the limitations of this study, chatbots showed similar overall accuracy and performed better on earlier exams.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Comparison of chatbots’ accuracy in endodontics questions in dentistry specialization exam in Türkiye: ChatGPT-4o, Gemini Advanced, Copilot, and Claude — 科研速览 Science Skim