Satya Ranjan Misra, Rupsa Das
Recent claims of expert-level endodontic diagnostic performance by large language models require cautious interpretation. In a 40-case, text-only comparison, methodological concerns limit reproducibility and clinical generalisability. The reported testing period preceded the public release of GPT-5, making clarification of the exact model version, platform, access dates, and settings essential. Use of a single endodontist as the reference standard demonstrates agreement with one clinician rather than independently verified diagnostic accuracy. Binary yes/no vignettes may also simplify the diagnostic task and do not reflect routine multimodal assessment incorporating radiographic interpretation. Single-query testing prevents evaluation of response consistency, while perfect accuracy in a small sample remains compatible with considerable statistical uncertainty. These findings are therefore best viewed as preliminary evidence of high concordance under controlled conditions. Future studies should use transparent model reporting, repeated runs, consensus reference standards, and clinically representative multimodal cases before expert-level performance is inferred.