科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Bioengineering2026-01-17· United States Medical Licensing Examination

Large Language Models Evaluation of Medical Licensing Examination Using GPT-4.0, ERNIE Bot 4.0, and GPT-4o

Luoyu Lian, Xin Luo, Kavimbi Chipusu, Muhammad Ashraf, Kelvin K. L. Wong, Wenjun Zhang

原始摘要(英文原文)· Original abstract
This study systematically evaluated the performance of three advanced large language models (LLMs)-GPT-4.0, ERNIE Bot 4.0, and GPT-4o-in the 2023 Chinese Medical Licensing Examination. Employing a dataset of 600 standardized questions, we analyzed the accuracy of each model in answering questions from three comprehensive sections: Basic Medical Comprehensive, Clinical Medical Comprehensive, and Humanities and Preventive Medicine Comprehensive. Our results demonstrate that both ERNIE Bot 4.0 and GPT-4o significantly outperformed GPT-4.0, achieving accuracies above the national pass mark. The study further examined the strengths and limitations of each model, providing insights into their applicability in medical education and potential areas for future improvement. These findings underscore the promise and challenges of deploying LLMs in multilingual medical education, suggesting a pathway towards integrating AI into medical training and assessment practices.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Large Language Models Evaluation of Medical Licensing Examination Using GPT-4.0, ERNIE Bot 4.0, and GPT-4o — 科研速览 Science Skim