Jasmine Mary Antony, Nikhil Harikrishnan, Nishi Jayasheelan
All three LLMs demonstrated high accuracy on operative dentistry MCQs, suggesting their potential utility as a supplementary educational tool. However, the residual error rates in the answers warrant the cautious integration of LLMs into dental education.
INTRODUCTION: Artificial intelligence (AI) has emerged as a valuable tool in the field of dentistry to assist healthcare providers in diagnosis, optimize treatment outcomes, support research, and improve education. Furthermore, the generative aspects of AI have emerged as a powerful tool for dental professionals to tackle clinical and academic challenges. However, the reliability of AI in domains such as undergraduate training remains unverified.
AIM: This study aimed to evaluate the accuracy and consistency of three large language models (LLMs) in answering multiple-choice questions on the undergraduate operative dentistry curriculum to determine their reliability as a supplementary learning resource.
METHODS: Sixty multiple-choice questions were formulated from the undergraduate operative dentistry curriculum. Each LLM (Claude, Gemini, and ChatGPT) was queried individually nine times over three days. The results were evaluated using a predetermined answer key. The accuracy of the LLMs was compared using a generalized estimating equation (GEE) with question-level clustering analysis. Consistency was assessed at the question level using response consistency, correctness consistency, and Fleiss' kappa. Statistical significance was set at P < 0.05.
RESULTS: Gemini demonstrated the highest overall accuracy (98.89%), followed by ChatGPT (98.33%) and Claude (97.78%). No statistically significant difference in accuracy was observed among the three LLMs (GEE, overall Wald χ 2(2) = 1.46, p = 0.481). All three models showed almost perfect intersession agreement (Fleiss' κ = 0.93-0.95), with response and correctness consistency ranging from 88.3% to 91.7% across models.
CONCLUSIONS: All three LLMs demonstrated high accuracy on operative dentistry MCQs, suggesting their potential utility as a supplementary educational tool. However, the residual error rates in the answers warrant the cautious integration of LLMs into dental education.