M K Nasreen, Paul Chalakkal, Harsh Ulhas Manerkar, Ramya Ramanathan, Mira Miranda, Saloni Prabhakar Pednekar
GP was found to have the highest accuracy for all parameters among the LLMs tested. However, all the LLMs were found to be inferior to expert human judgment.
BACKGROUND: Various studies have evaluated the accuracy of large language models (LLMs) in the detection of periapical radiolucencies from adult radiographs and orthopantomographs. However, none have evaluated their efficacy in evaluating parameters from mixed dentition radiographs.
AIM: The aim of this study was to evaluate the accuracy of GPT-5 (GT), Gemini 3 Pro (GP), and Sonar (SR) in assessing tooth detection and identification, dental caries, periapical status, and stages of tooth development.
SETTINGS AND DESIGN: A total of 93 mixed dentition radiographs from children aged 6-12 years were included. Two pediatric dentists evaluated anonymized radiographs and were blinded to the outputs generated by the LLM.
METHODOLOGY: A standardized prompt-engineering framework was used to instruct the LLM. The outputs generated were recorded in a Microsoft Excel spreadsheet and compared with the reference standard established by the examiners.
STATISTICAL ANALYSIS USED: Stata Statistical Software (Release 19), Cohen's kappa coefficient, Cochran's Q test, and McNemar's test were used for statistical analysis. Statistical significance was established at P < 0.05.
RESULTS: GP was found to have the highest accuracy in identifying teeth. GP showed the highest observed and expected agreements in identifying caries levels. GP was better at recognizing advanced carious lesions, while GT and SR were better at identifying sound teeth. None of the LLMs could identify periodontitis. GP showed the best overall performance in identifying tooth development.
CONCLUSIONS: GP was found to have the highest accuracy for all parameters among the LLMs tested. However, all the LLMs were found to be inferior to expert human judgment.