Pingrong Lin, Qin Deng, Ying Zhou
As large language models (LLMs) become deeply integrated into the educational landscape, evaluation criteria focusing solely on performance are insufficient to mitigate the risks of value misalignment and socioethical concerns. To steer educational LLMs towards responsible and beneficial development, this study aims to construct a multidimensional evaluation framework grounded in educational theory. Initially, a preliminary pool of evaluation indicators was established on the basis of a review of the literature and pedagogical theories. The Delphi method was subsequently employed to refine the indicator structure by integrating opinions from 21 cross-disciplinary experts. The analytic hierarchy process (AHP) was then applied to weigh these indicators and determine their priorities. The final framework comprises 5 first-level indicators and 21 s-level indicators. Learning effectiveness, knowledge construction capability, and social alignment are assigned critical weights, whereas intelligent interaction capability is less prioritized. Among the second-level indicators, information veracity was weighted the highest, while educational equity had the weakest influence. This study not only provides direction for the development and optimization of educational LLMs but also offers a reference for establishing a responsible artificial intelligence in education ecosystem.