科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Odontology2026-09-08

Large language models answering structured questions on external cervical resorption: a comparative study.

Alexandre Barbosa, Ana Cristina Braga, Irene Pina-Vaz, Inês Ferreira

原始摘要(英文原文)· Original abstract
This study aimed to evaluate the accuracy, consistency, and temporal stability of six large language models (LLMs), when answering structured questions on external cervical resorption (ECR) across seven consecutive days, while assessing domain-specific error patterns. Four Base configurations (ChatGPT-5, Gemini, Claude, and Mistral) and two document-grounded (RAG) configurations (NotebookLM and Perplexity Pro) were evaluated using 46 validated dichotomous questions on ECR. Each model was queried using three independent user accounts over seven consecutive days. Accuracy was defined as agreement with predefined gold standard answers. Inter-account consistency and temporal stability were evaluated across accounts and days. Domain-specific error patterns were analyzed. Generalized estimating equation models were used to evaluate differences in response accuracy across LLM configurations, clinical domains, user accounts, and evaluation days. Overall accuracy was high (90.8%). Model configuration significantly influenced performance (p = 0.050), with NotebookLM and Gemini showing the lowest estimated error rates (3.0% and 4.0%, respectively), while Claude demonstrated the highest error rate (12.0%). Accuracy remained stable across user accounts and evaluation days, with no evidence of account-related variability (p = 0.654) or temporal drift (p = 0.875). Clinical domain significantly affected performance (p < 0.001); classification questions showed the highest error rate (37.0%). LLMs demonstrated high accuracy, reproducibility, and temporal stability. However, performance varied across models and clinical domains, with classification questions remaining challenging. These findings support the cautious integration of LLMs into clinical practice and endodontic education while underscoring the need for human oversight. Importantly, observed performance was restricted to standardized questions and should not be interpreted as evidence of comprehensive clinical decision-making capability.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Large language models answering structured questions on external cervical resorption: a comparative study. — 科研速览 Science Skim