科研速览继续刷下去 →
◆ Frontiers in dental medicine2026-01-01

Accuracy and response repeatability of three large language models on undergraduate operative dentistry multiple-choice questions.

Jasmine Mary Antony, Nikhil Harikrishnan, Nishi Jayasheelan

一句话结论

All three LLMs demonstrated high accuracy on operative dentistry MCQs, suggesting their potential utility as a supplementary educational tool. However, the residual error rates in the answers warrant the cautious integration of LLMs into dental education.

原始摘要(原文)
INTRODUCTION: Artificial intelligence (AI) has emerged as a valuable tool in the field of dentistry to assist healthcare providers in diagnosis, optimize treatment outcomes, support research, and improve education. Furthermore, the generative aspects of AI have emerged as a powerful tool for dental professionals to tackle clinical and academic challenges. However, the reliability of AI in domains such as undergraduate training remains unverified. AIM: This study aimed to evaluate the accuracy and consistency of three large language models (LLMs) in answering multiple-choice questions on the undergraduate operative dentistry curriculum to determine their reliability as a supplementary learning resource. METHODS: Sixty multiple-choice questions were formulated from the undergraduate operative dentistry curriculum. Each LLM (Claude, Gemini, and ChatGPT) was queried individually nine times over three days. The results were evaluated using a predetermined answer key. The accuracy of the LLMs was compared using a generalized estimating equation (GEE) with question-level clustering analysis. Consistency was assessed at the question level using response consistency, correctness consistency, and Fleiss' kappa. Statistical significance was set at P < 0.05. RESULTS: Gemini demonstrated the highest overall accuracy (98.89%), followed by ChatGPT (98.33%) and Claude (97.78%). No statistically significant difference in accuracy was observed among the three LLMs (GEE, overall Wald χ 2(2) = 1.46, p = 0.481). All three models showed almost perfect intersession agreement (Fleiss' κ = 0.93-0.95), with response and correctness consistency ranging from 88.3% to 91.7% across models. CONCLUSIONS: All three LLMs demonstrated high accuracy on operative dentistry MCQs, suggesting their potential utility as a supplementary educational tool. However, the residual error rates in the answers warrant the cautious integration of LLMs into dental education.
读原文 ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文

Accuracy and response repeatability of three large language models on undergraduate operative dentistry multiple-choice questions. — 科研速览 Science Skim