科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Journal of dental sciences2026-01-01

Artificial intelligence-powered chatbots' responses to orthodontic questions from the dentistry specialization examination: Accuracy and source evaluation.

Berrak Çakmak, Tevhide Sökmen, Burcu Baloş Tuncer

一句话结论 · In one sentence

Chatbots exhibited strong text-based reasoning abilities but limited visual interpretation skills. While ChatGPT-5.0 provided more reliable and well-referenced responses, other models showed weaker citation practices. These underscored both the potential and the current limitations of AI-based systems in orthodontic education and clinical practice.

原始摘要(英文原文)· Original abstract
BACKGROUND: /purpose: The use of artificial intelligence (AI) powered chatbots in dental education is becoming increasingly widespread. Evaluating their performance and the reliability of their sources is essential to understand their educational value. The aim of this study was to evaluate the performance of AI-powered chatbots in addressing orthodontic questions from the Dental Specialty Exam (DUS) and to assess the accuracy and reliability of the information sources on which they rely. MATERIALS AND METHODS: A total of 129 orthodontic questions from the exam administered between 2012 and 2021 were categorized according to Bloom's taxonomy. Each question was individually entered into ChatGPT-5, Claude 3.7, and Copilot, and their performances were comparatively evaluated. The sources referenced by the chatbots while generating their answers were also assessed. The data were analyzed using Pearson's chi-squared test. RESULTS: ChatGPT-5, Claude 3.7, and Copilot achieved accuracy rates of 82.2 %, 83.7 %, and 85.3 %, respectively. Copilot performed best on scenario-based questions (100 %) but performed worst on visual analysis questions (33.3 %). Citation analysis showed that, ChatGPT-5.0 used reliable academic sources, whereas Claude cited few and less credible references, and Copilot relied mainly on moderately reliable materials. CONCLUSION: Chatbots exhibited strong text-based reasoning abilities but limited visual interpretation skills. While ChatGPT-5.0 provided more reliable and well-referenced responses, other models showed weaker citation practices. These underscored both the potential and the current limitations of AI-based systems in orthodontic education and clinical practice.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Artificial intelligence-powered chatbots' responses to orthodontic questions from the dentistry specialization examination: Accuracy and source evaluation. — 科研速览 Science Skim