科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Surgical and radiologic anatomy : SRA2026-08-31

Comparison of artificial intelligence and lecturers' grading in the evaluation of handwritten open-ended anatomy examinations.

Mert Nahir, Abdulkerim Kasap, Canan Bayraktar Nahir

一句话结论 · In one sentence

AI chatbot grading performance varied across models. ChatGPT produced scores most closely aligned with lecturers' grading, whereas providing a standardized answer key did not consistently improve performance. AI-assisted grading may support assessment in anatomy education, but human oversight and structured grading criteria remain essential, particularly in high-stakes examinations.

原始摘要(英文原文)· Original abstract
PURPOSE: This study evaluated the performance of artificial intelligence (AI) chatbots in grading handwritten open-ended anatomy examinations and examined whether providing a standardized answer key improved agreement with lecturers' grading. METHODS: Examination papers from 81 s-year dental students were included. The examination comprised 10 open-ended questions, each worth 10 points. Two anatomists graded all papers using a standardized answer key, and the mean of their scores was used as lecturers' grading. Four AI chatbots-ChatGPT, Gemini, Copilot, and DeepSeek-graded each paper under two conditions: without an answer key and with an answer key and detailed grading criteria. Agreement with lecturers' grading was assessed using the intraclass correlation coefficient (ICC). RESULTS: Differences were observed among the nine grading methods (p < 0.001). No significant differences were found between lecturers' grading and ChatGPT scores obtained either without or with the answer key. ChatGPT showed the highest agreement with lecturers' grading after the answer key was provided (ICC = 0.923, 95% CI [0.883, 0.950]), followed by ChatGPT without the answer key (ICC = 0.904, 95% CI [0.855, 0.937]). Gemini and Copilot showed good agreement without the answer key, but agreement decreased when the answer key was provided. DeepSeek demonstrated the lowest agreement under both conditions. CONCLUSION: AI chatbot grading performance varied across models. ChatGPT produced scores most closely aligned with lecturers' grading, whereas providing a standardized answer key did not consistently improve performance. AI-assisted grading may support assessment in anatomy education, but human oversight and structured grading criteria remain essential, particularly in high-stakes examinations.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Comparison of artificial intelligence and lecturers' grading in the evaluation of handwritten open-ended anatomy examinations. — 科研速览 Science Skim