科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Medical science monitor : international medical journal of experimental and clinical research2026-09-04

Accuracy, Self-Reported Confidence, and Overconfidence of Large Language Models in Endodontics: An Evaluation Using National Specialty Examination Questions.

Mehmet Adıgüzel

原始摘要(英文原文)· Original abstract
BACKGROUND This study aimed to evaluate the accuracy, confidence, and overconfidence behavior of large language models (LLMs) in endodontics using questions derived from a national specialty entrance examination. MATERIAL AND METHODS A total of 123 text-based endodontic questions from the Turkish Dental Specialty Examination (2017-2026) were included after excluding annulled and image-based questions. Three LLMs (ChatGPT, Claude, and Gemini) were assessed. Each model answered all questions using a standardized prompt and provided a confidence score (0%-100%). Accuracy was recorded as correct/incorrect. Overconfidence was defined as incorrect responses with ≥90% confidence. Statistical analyses were performed using Cochran's Q test, Friedman test, and post hoc pairwise comparisons with Bonferroni correction. RESULTS Accuracy rates were 89.4% for ChatGPT, 76.4% for Claude, and 90.2% for Gemini, with significant differences among models (χ²(2)=21.00, P<0.001). Confidence scores differed significantly (c2(2)=213.40, P<0.001), with Gemini demonstrating highest confidence (99.51±1.49), followed by ChatGPT (90.88±9.19) and Claude (80.22±11.10). Overconfidence rates were 9.8% for Gemini and 5.7% for ChatGPT, while no overconfident responses were observed for Claude (χ²(2)=14.53, P=0.001). Despite similar accuracy between ChatGPT and Gemini, confidence patterns differed markedly, demonstrating that comparable accuracy does not necessarily reflect comparable reliability. CONCLUSIONS LLMs demonstrated high accuracy in answering text-based endodontic examination questions; however, significant differences were observed in confidence behavior and overconfidence patterns. The presence of high-confidence incorrect responses suggests that accuracy alone may be insufficient to fully evaluate model reliability. These findings highlight the importance of considering confidence-related behavior alongside accuracy when assessing LLM performance in examination-style endodontic tasks.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Accuracy, Self-Reported Confidence, and Overconfidence of Large Language Models in Endodontics: An Evaluation Using National Specialty Examination Questions. — 科研速览 Science Skim