William Villegas-Ch, Aracely Mera-Navarrete, Fernando Zúñiga-Tello, Rommel Gutiérrez
The review identified two broad trajectories. The first is instructional and developmental, where generative AI, AI tutoring, and AI-generated materials are used or proposed to support differentiation, writing, questioning, creativity, mentoring, and personalized learning. The second is assessment-oriented, where machine learning and related models are applied to gifted and talented identification, twice-exceptional identification, and decision support. Across both trajectories, the literature positions AI less as a replacement for educators than as a tool whose value depends on task design, teacher judgment, and human oversight. Reported opportunities include support for differentiated preparation, advanced learning, creative production, and the recognition of complex learner profiles. The reported risks include bias in identification data, overreliance, threats to originality, privacy concerns, limited transparency, and insufficient preparation among educators and institutions.
Large language models (LLMs) are increasingly used to evaluate open-ended educational responses. However, their performance is often assessed using aggregate metrics that provide limited insight into prediction stability, uncertainty, error patterns, and feedback quality. This study presents EduFairBench, a reproducible evaluation protocol designed to characterize LLM behavior across short-answer assessment and automated essay scoring using open educational benchmarks. The protocol combines repeated inference, majority-vote consolidation, uncertainty estimation, error analysis, and structural evaluation of generated feedback within a unified experimental framework. Experiments were conducted on SciEntsBank, Beetle, and ASAP2, comprising 2,000 student responses and 10,000 independent LLM inferences. The results showed moderate predictive agreement with human assessment while revealing substantial differences between nominal and ordinal evaluation tasks. Repeated inference demonstrated high internal stability across benchmarks, although systematic errors remained in semantically adjacent categories, indicating that prediction consistency does not necessarily imply correctness. Feedback quality varied by task type, with longer textual contexts yielding more specific and pedagogically structured explanations. These findings demonstrate that evaluating educational LLMs requires complementary analyses beyond conventional performance metrics. EduFairBench provides a reproducible methodology for jointly analyzing predictive performance, robustness, uncertainty, and feedback quality, providing a comprehensive methodological framework for the rigorous evaluation of LLM-based educational assessment systems.