Hikaru Watanabe, Osamu Uehara, Tetsuro Morikawa, Takayuki Kojima, Takayuki Suga, Akira Toyofuku, Satoshi Takada, Yoshihiro Abiko
LLMs demonstrated moderate diagnostic performance for image-based dental pathology questions, with Gemini demonstrating superior accuracy and consistency. Although promising decision support tools in education and clinical settings, LLMs still exhibit domain-specific limitations and require careful oversight. The integration of explainable artificial intelligence and real-world clinical validation is recommended for its safe and effective use.
BACKGROUND /PURPOSE: Large language models (LLMs), such as Chat Generative Pre-trained Transformer (ChatGPT) and Gemini, have demonstrated promising capabilities for medical question-answering tasks. However, the diagnostic performance of LLMs in image-based oral pathologies remains largely unexplored. This study aimed to evaluate these capabilities using histopathological images obtained from the Japanese National Dental Examination.
MATERIALS AND METHODS: This study aimed to evaluate and compare the diagnostic accuracy and agreement of three LLMs (ChatGPT-4o [ChatGPT], Gemini 1.5 Pro [Gemini], and Claude 3.5 Sonnet [Claude]) on pathology image-based questions from the Japanese National Dental Examination.
RESULTS: Gemini achieved the highest accuracy (61.4 %), followed by Claude (52.3 %), and ChatGPT (45.4 %). Gemini and ChatGPT exhibited significant differences (P = 0.00054). Cohen's kappa values indicated moderate agreement for all models, with Gemini showing the highest agreement (κ = 0.599). Accuracy varied across disease categories: Gemini excelled in squamous cell carcinoma (92.0 %) and salivary gland tumors, whereas Claude performed best on soft tissue lesions. The confusion matrix analysis revealed distinct misclassification patterns in each model, particularly between odontogenic tumors and cystic lesions.
CONCLUSION: LLMs demonstrated moderate diagnostic performance for image-based dental pathology questions, with Gemini demonstrating superior accuracy and consistency. Although promising decision support tools in education and clinical settings, LLMs still exhibit domain-specific limitations and require careful oversight. The integration of explainable artificial intelligence and real-world clinical validation is recommended for its safe and effective use.