科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Journal of dental sciences2026-01-01

Performance of large language models on image-based oral pathology questions from the Japanese National Dental Examination.

Hikaru Watanabe, Osamu Uehara, Tetsuro Morikawa, Takayuki Kojima, Takayuki Suga, Akira Toyofuku, Satoshi Takada, Yoshihiro Abiko

一句话结论 · In one sentence

LLMs demonstrated moderate diagnostic performance for image-based dental pathology questions, with Gemini demonstrating superior accuracy and consistency. Although promising decision support tools in education and clinical settings, LLMs still exhibit domain-specific limitations and require careful oversight. The integration of explainable artificial intelligence and real-world clinical validation is recommended for its safe and effective use.

原始摘要(英文原文)· Original abstract
BACKGROUND /PURPOSE: Large language models (LLMs), such as Chat Generative Pre-trained Transformer (ChatGPT) and Gemini, have demonstrated promising capabilities for medical question-answering tasks. However, the diagnostic performance of LLMs in image-based oral pathologies remains largely unexplored. This study aimed to evaluate these capabilities using histopathological images obtained from the Japanese National Dental Examination. MATERIALS AND METHODS: This study aimed to evaluate and compare the diagnostic accuracy and agreement of three LLMs (ChatGPT-4o [ChatGPT], Gemini 1.5 Pro [Gemini], and Claude 3.5 Sonnet [Claude]) on pathology image-based questions from the Japanese National Dental Examination. RESULTS: Gemini achieved the highest accuracy (61.4 %), followed by Claude (52.3 %), and ChatGPT (45.4 %). Gemini and ChatGPT exhibited significant differences (P = 0.00054). Cohen's kappa values indicated moderate agreement for all models, with Gemini showing the highest agreement (κ = 0.599). Accuracy varied across disease categories: Gemini excelled in squamous cell carcinoma (92.0 %) and salivary gland tumors, whereas Claude performed best on soft tissue lesions. The confusion matrix analysis revealed distinct misclassification patterns in each model, particularly between odontogenic tumors and cystic lesions. CONCLUSION: LLMs demonstrated moderate diagnostic performance for image-based dental pathology questions, with Gemini demonstrating superior accuracy and consistency. Although promising decision support tools in education and clinical settings, LLMs still exhibit domain-specific limitations and require careful oversight. The integration of explainable artificial intelligence and real-world clinical validation is recommended for its safe and effective use.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Performance of large language models on image-based oral pathology questions from the Japanese National Dental Examination. — 科研速览 Science Skim