Büşra Karaca, Yunus Emre Çakmak, Damla Erkal
Large language models (LLMs) are increasingly used in healthcare, but their performance in endodontic decision-making remains unclear. This study aimed to compare six LLMs in terms of diagnostic appropriateness for endodontic treatment planning. Fifty clinical scenarios were developed and entered into six LLMs (ChatGPT-4o, ChatGPT-3.5, Claude 4, Copilot, DeepSeek-V3, Gemini 2.5). Two specialists scored responses as appropriate or inappropriate. Repeated measures ANOVA and chi-square tests were used for analysis. Claude showed the highest accuracy (76%), followed by DeepSeek and Gemini. ChatGPT-3.5 had the lowest (40%). Significant differences were found between models (p < 0.05). Performance was better on straightforward cases than on complex scenarios. LLMs vary widely in diagnostic accuracy for endodontic cases. While some models show promise, others may provide confidently incorrect recommendations. Caution and human oversight remain essential until domain-specific, fine-tuned models are developed.