Mert Nahir, Abdulkerim Kasap, Canan Bayraktar Nahir
AI chatbot grading performance varied across models. ChatGPT produced scores most closely aligned with lecturers' grading, whereas providing a standardized answer key did not consistently improve performance. AI-assisted grading may support assessment in anatomy education, but human oversight and structured grading criteria remain essential, particularly in high-stakes examinations.
PURPOSE: This study evaluated the performance of artificial intelligence (AI) chatbots in grading handwritten open-ended anatomy examinations and examined whether providing a standardized answer key improved agreement with lecturers' grading.
METHODS: Examination papers from 81 s-year dental students were included. The examination comprised 10 open-ended questions, each worth 10 points. Two anatomists graded all papers using a standardized answer key, and the mean of their scores was used as lecturers' grading. Four AI chatbots-ChatGPT, Gemini, Copilot, and DeepSeek-graded each paper under two conditions: without an answer key and with an answer key and detailed grading criteria. Agreement with lecturers' grading was assessed using the intraclass correlation coefficient (ICC).
RESULTS: Differences were observed among the nine grading methods (p < 0.001). No significant differences were found between lecturers' grading and ChatGPT scores obtained either without or with the answer key. ChatGPT showed the highest agreement with lecturers' grading after the answer key was provided (ICC = 0.923, 95% CI [0.883, 0.950]), followed by ChatGPT without the answer key (ICC = 0.904, 95% CI [0.855, 0.937]). Gemini and Copilot showed good agreement without the answer key, but agreement decreased when the answer key was provided. DeepSeek demonstrated the lowest agreement under both conditions.
CONCLUSION: AI chatbot grading performance varied across models. ChatGPT produced scores most closely aligned with lecturers' grading, whereas providing a standardized answer key did not consistently improve performance. AI-assisted grading may support assessment in anatomy education, but human oversight and structured grading criteria remain essential, particularly in high-stakes examinations.