Shuma Hamaguchi, Masakazu Hamada, Shunya Ikeda, Satoru Kusaka, Tatsuya Akitomo, Ryota Nomura
This report suggests that generative AI is improving annually and adapting to the National Dental Examination. Although each AI model is suited to different fields, the trends may change over time. It is necessary to continue comparing and analyzing AI models and provide users with the latest information.
BACKGROUND/PURPOSE: Artificial intelligence (AI) has become widely used and applied in various fields. Although several studies have been conducted using generative AI for various qualification exams, to the best of our knowledge, none has focused on performance changes over time.
MATERIALS AND METHODS: In August 2025, ChatGPT 5, Gemini 2.5, Microsoft Copilot, and Medi-Search were asked to answer compulsory questions from five years of the Japanese National Dental Examination. In 2024, we also conducted similar tests on other ChatGPT series and Gemini, and the scores were compared.
RESULTS: In 2025, Copilot, Gemini, and MediSearch scored 80 % or higher, which was the passing standard, for all five years. Although ChatGPT 3.5 did not meet the passing standard for any of the five years, ChatGPT 4o mini and ChatGPT 5 exceeded it for two and three years, respectively. In addition, both Chat GPT's and Gemini's scores substantially improved over time and with each update.
CONCLUSION: This report suggests that generative AI is improving annually and adapting to the National Dental Examination. Although each AI model is suited to different fields, the trends may change over time. It is necessary to continue comparing and analyzing AI models and provide users with the latest information.