科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Computers and Education Open2026-04-20· Grading (engineering)

A systematic comparison of Large Language Models for automated assignment assessment in programming education: Exploring the importance of architecture and vendor

Marcin Jukiewicz

原始摘要(英文原文)· Original abstract
This study presents the first large-scale, side-by-side comparison of contemporary Large Language Models (LLMs) in the automated grading of programming assignments. Drawing on over 6000 student submissions collected across four years of an introductory programming course, the study systematically analyzed grade distributions, differences in mean scores and variability reflecting stricter or more lenient grading, and the consistency and clustering of grading patterns across models. Eighteen publicly available models were evaluated: Anthropic (claude-3-5-haiku, claude-opus-4-1, claude-sonnet-4), DeepSeek (deepseek-chat, deepseek-reasoner), Google (gemini-2.0-flash-lite, gemini-2.0-flash, gemini-2.5-flash-lite, gemini-2.5-flash, gemini-2.5-pro), and OpenAI (gpt-4.1-mini, gpt-4.1-nano, gpt-4.1, gpt-4o-mini, gpt-4o, gpt-5-mini, gpt-5-nano, gpt-5). Statistical analyses, including correlation, agreement, and clustering methods, revealed clear and systematic differences in grading behavior across models. Distinct grading patterns emerged, ranging from more lenient to more restrictive evaluation styles, while models from the same vendor tended to cluster together, suggesting shared algorithmic approaches to code assessment. Full-scale models consistently outperformed their smaller “mini” and “nano” counterparts. Despite strong internal agreement among models, alignment with human teachers’ grades remained limited: even the best-performing model achieved only moderate reliability. These findings indicate that the choice of LLM for educational deployment is not neutral and may substantially influence grading outcomes. The results highlight the importance of careful model selection, transparent reporting of evaluation metrics, and a human-in-the-loop approach when integrating AI-based grading systems in educational contexts.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

A systematic comparison of Large Language Models for automated assignment assessment in programming education: Exploring the importance of architecture and vendor — 科研速览 Science Skim