科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Medical Science Educator2026-03-03· Rubric

Can Generative Artificial Intelligence Reliably Score Open-Ended Question Assessments in Undergraduate Medical Education?

Doreen M. Olvet, Marieke Kruidering, Tracy B. Fulton, Bao C. Q. Truong, Kumiko Endo, Robert Lucito, Joanne M. Willey

原始摘要(英文原文)· Original abstract
Abstract There are numerous benefits to including open-ended questions (OEQs) in the assessment of medical knowledge, but one of the biggest challenges is the time it takes to grade student responses. With the widescale introduction of generative artificial intelligence (AI), it is plausible that OEQ exams can be automatically scored. The purpose of this study was to establish the accuracy of generative AI when scoring medical student OEQ exams. Students’ responses from OEQs administered at two US allopathic medical schools were analyzed. Case vignettes, questions, rubrics, and student responses were fed into the GPT-4 model via the Med2Lab platform. The Med2Lab system was specifically engineered to manage rubric integration and automate prompt workflows. Scores and feedback on students’ responses were generated and compared to faculty scores using Cohen’s weighted kappa (k w ) to evaluate inter-rater reliability (IRR). An error pattern analysis was performed to assess why there were scoring discrepancies between faculty and GPT-4, then this information was used to perform rubric engineering. We ran 3 iterations of GPT-4 scoring after each rubric adjustment. By the third iteration, IRR between faculty and GPT-4 was substantial using the analytic rubric (question 1A: k w =0.94; question 2A: k w =0.88) and the holistic rubric (question 2H: k w =0.89). IRR for question 1H reached moderate reliability (k w =0.54). We identified errors in GPT-4 and faculty scoring, although score discrepancies were typically only 1-point. Our data suggest that generative AI can be used to reliably score OEQ exams using an iterative process of rubric engineering to achieve maximum reliability.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Can Generative Artificial Intelligence Reliably Score Open-Ended Question Assessments in Undergraduate Medical Education? — 科研速览 Science Skim