科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Australian Endodontic Journal2025-11-21· Relevance (law)

Clinical Relevance of Large Language Models in Endodontics: Diagnostic Appropriateness Based on 50 Simulated Case Scenarios

Büşra Karaca, Yunus Emre Çakmak, Damla Erkal

原始摘要(英文原文)· Original abstract
Large language models (LLMs) are increasingly used in healthcare, but their performance in endodontic decision-making remains unclear. This study aimed to compare six LLMs in terms of diagnostic appropriateness for endodontic treatment planning. Fifty clinical scenarios were developed and entered into six LLMs (ChatGPT-4o, ChatGPT-3.5, Claude 4, Copilot, DeepSeek-V3, Gemini 2.5). Two specialists scored responses as appropriate or inappropriate. Repeated measures ANOVA and chi-square tests were used for analysis. Claude showed the highest accuracy (76%), followed by DeepSeek and Gemini. ChatGPT-3.5 had the lowest (40%). Significant differences were found between models (p < 0.05). Performance was better on straightforward cases than on complex scenarios. LLMs vary widely in diagnostic accuracy for endodontic cases. While some models show promise, others may provide confidently incorrect recommendations. Caution and human oversight remain essential until domain-specific, fine-tuned models are developed.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Clinical Relevance of Large Language Models in Endodontics: Diagnostic Appropriateness Based on 50 Simulated Case Scenarios — 科研速览 Science Skim