科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Journal of dental sciences2026-01-01

Performance of ChatGPT-4, Gemini, and DeepSeek-V3 on English-translated questions from the Taiwan national dental Technician Licensing examination over a three-week period.

Yi-Pang Lee, Ching-Yi Huang, Andy Sun, Chun-Pin Chiang

一句话结论 · In one sentence

ChatGPT-4, Gemini, and DeepSeek-V3 demonstrate moderate capability in answering dental technician examination questions but generally show no significant improvement over a three-week period. Translating Chinese-language questions into English may improve the performance of ChatGPT-4 and Gemini.

原始摘要(英文原文)· Original abstract
BACKGROUND /PURPOSE: Large language models (LLMs) have shown potential in answering professional examination questions. This study evaluated the performance of ChatGPT-4, Gemini, and DeepSeek-V3 in answering English-translated questions from the 2023 Taiwan National Dental Technician Licensing Examination (TNDTLE) over a three-week period. MATERIALS AND METHODS: A total of 194 English-translated, text-based multiple-choice questions were selected from the 2023 TNDTLE. ChatGPT-4, Gemini, and DeepSeek-V3 were used to answer the same set of English-translated questions at four time points: baseline and one-, two-, and three-week follow-ups. Accuracy rates (ARs) were calculated and compared to evaluate changes over time and differences among the three LLMs, between basic and clinical subjects, and between English-translated and original Chinese-language questions. RESULTS: The baseline ARs were 69.1% for ChatGPT-4, 75.8% for Gemini, and 69.6% for DeepSeek- V3. Among the three LLMs, only ChatGPT-4 demonstrated a statistically significant improvement at the three-week follow-up (P = 0.029). No significant differences were observed among the three LLMs at most time points, except that Gemini achieved a significantly higher AR than Deep-Seek-V3 at the three-week follow-up (78.4% vs. 68.6%, P = 0.039). ARs were generally higher for basic subjects than for clinical subjects. ChatGPT-4 and Gemini achieved significantly higher ARs for English-translated questions than for original Chinese-language questions, whereas DeepSeek- V3 showed no significant language-related difference. CONCLUSION: ChatGPT-4, Gemini, and DeepSeek-V3 demonstrate moderate capability in answering dental technician examination questions but generally show no significant improvement over a three-week period. Translating Chinese-language questions into English may improve the performance of ChatGPT-4 and Gemini.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Performance of ChatGPT-4, Gemini, and DeepSeek-V3 on English-translated questions from the Taiwan national dental Technician Licensing examination over a three-week period. — 科研速览 Science Skim