Alexandre Barbosa, Ana Cristina Braga, Irene Pina-Vaz, Inês Ferreira
This study aimed to evaluate the accuracy, consistency, and temporal stability of six large language models (LLMs), when answering structured questions on external cervical resorption (ECR) across seven consecutive days, while assessing domain-specific error patterns. Four Base configurations (ChatGPT-5, Gemini, Claude, and Mistral) and two document-grounded (RAG) configurations (NotebookLM and Perplexity Pro) were evaluated using 46 validated dichotomous questions on ECR. Each model was queried using three independent user accounts over seven consecutive days. Accuracy was defined as agreement with predefined gold standard answers. Inter-account consistency and temporal stability were evaluated across accounts and days. Domain-specific error patterns were analyzed. Generalized estimating equation models were used to evaluate differences in response accuracy across LLM configurations, clinical domains, user accounts, and evaluation days. Overall accuracy was high (90.8%). Model configuration significantly influenced performance (p = 0.050), with NotebookLM and Gemini showing the lowest estimated error rates (3.0% and 4.0%, respectively), while Claude demonstrated the highest error rate (12.0%). Accuracy remained stable across user accounts and evaluation days, with no evidence of account-related variability (p = 0.654) or temporal drift (p = 0.875). Clinical domain significantly affected performance (p < 0.001); classification questions showed the highest error rate (37.0%). LLMs demonstrated high accuracy, reproducibility, and temporal stability. However, performance varied across models and clinical domains, with classification questions remaining challenging. These findings support the cautious integration of LLMs into clinical practice and endodontic education while underscoring the need for human oversight. Importantly, observed performance was restricted to standardized questions and should not be interpreted as evidence of comprehensive clinical decision-making capability.