Wentong Yang, Ting Xu, Junjie Wei, Wentao Zheng, Wenbo Yan, Jingkai Wang, Guangdi Chu, Haitao Niu
DeepSeek-assisted self-study was associated with higher posttest scores than traditional internet-based learning, whereas ChatGPT showed numerically higher but nonsignificant results. These findings offer evidence-based insights into the embedding of GenAI within medical education frameworks, providing guidance for educators in developing teaching strategies and for institutions in formulating relevant policies.
BACKGROUND: Since its release in November 2022, generative AI (GenAI) tools, including ChatGPT, have gained widespread attention across various sectors, including medical education.
OBJECTIVE: This study seeks to examine the effectiveness and feasibility of GenAI tools (ChatGPT o3‑mini [OpenAI] and DeepSeek R1) in enhancing urology teaching outcomes for medical undergraduates.
METHODS: We assessed the accuracy of responses from ChatGPT o3-mini and DeepSeek R1 to authoritative urology multiple-choice questions. Then, a randomized controlled trial was performed to compare the learning outcomes of students using ChatGPT o3-mini and DeepSeek R1 with those using traditional learning methods. Additionally, a questionnaire was designed to survey medical undergraduates' perspectives on the application of AI in urology education.
RESULTS: DeepSeek R1 demonstrated higher accuracy than ChatGPT o3-mini in answering urology-related multiple-choice questions. In the test following the self-study period, the DeepSeek R1 group surpassed both the control and ChatGPT o3-mini groups in total scores across various question types. Despite the superior scores in the ChatGPT o3-mini group, statistical significance was not achieved relative to the control group. Survey results revealed that most students had a positive attitude toward AI-assisted learning, believing it could effectively enhance medical education.
CONCLUSIONS: DeepSeek-assisted self-study was associated with higher posttest scores than traditional internet-based learning, whereas ChatGPT showed numerically higher but nonsignificant results. These findings offer evidence-based insights into the embedding of GenAI within medical education frameworks, providing guidance for educators in developing teaching strategies and for institutions in formulating relevant policies.