科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Discover Artificial Intelligence2026-02-18· Generalizability theory

Exploring potential of large language models for automated essay scoring in education

Nimra Mughal, Ali Shariq Imran, Sher Muhammad Daudpota, Zenun Kastrati, Waheed Noor

原始摘要(英文原文)· Original abstract
Abstract The assessment of open-ended written work is of vital importance to the student learning experience. Conventional essay grading methods heavily depend on expert manual assessment, making them susceptible to errors due to fatigue, bias, and subjectivity. To address this, recent research has introduced AI-based Automated Essay Scoring (AES) systems. While most studies have concentrated on predicting scores, only a few have integrated AES systems with the well-known Large Language Models (LLMs). This study explores the application of LLMs, including GPT and Gemini for AES. The proposed approach was evaluated on two benchmark datasets, namely “Hewlett Foundation: Automated Essay Scoring (ASAP–AES)” and “Learning Agency Lab–Automated Essay Scoring 2.0 (LA–AES)”. The proposed method achieved promising results in AES, demonstrating effectiveness on both the benchmark datasets. Statistical analysis revealed that Gemini outperformed GPT, achieving an average Quadratic Weighted Kappa (QWK) score of 0.45 on the ASAP–AES and 0.43 on the LA–AES. To assess the generalizability and objectivity of the proposed approach, real-world data was collected from an O-Level classroom at Sukkur IBA Community College, Pakistan. Multiple human evaluators participated in the study to examine potential biases in human assessment. The findings indicate that LLM-based scoring demonstrates improved objectivity and reduced bias compared to human assessors.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Exploring potential of large language models for automated essay scoring in education — 科研速览 Science Skim