科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Frontiers in research metrics and analytics2026-01-01

Conditional validity in LLM-mediated L2 assessment: an argument-based systematic review and meta-analysis.

Latifah Hamdan Alghamdi, Talal Musaed Alghizzi

一句话结论 · In one sentence

Source-level auditing showed that directly comparable Pearson evidence was sparse. Three source-verified L2 studies yielded moderate human-LLM score correspondence (r = 0.66, 95% CI [0.53, 0.75]) with substantial heterogeneity (I 2 = 70.8%). Evidence was strongest in rubric-guided, structurally constrained, and human-supervised contexts, whereas construct representation, fairness, operational reproducibility, and generalizability remained underdeveloped. Across 14 studies, LLM-mediated formative feedback produced a small-to-moderate positive effect on revision quality and short-term writing outcomes (g = 0.42, 95% CI [0.27, 0.57]), although heterogeneity was substantial (I 2 = 69.8%) and concerns about cognitive offloading, learner agency, and reproducibility persisted.

原始摘要(英文原文)· Original abstract
INTRODUCTION: Large language models (LLMs) are increasingly used for scoring and feedback in second-language (L2) assessment, yet the validity of the resulting interpretations remains contested. This review evaluated when LLM-mediated assessment is psychometrically and educationally defensible using an argument-based validity framework focused on construct representation, reliability and reproducibility, fairness, and washback. METHODS: Following PRISMA 2020, we systematically searched nine databases and repositories for empirical studies published from January 2022 to December 2025. Fifty-two studies met the inclusion criteria. Quantitative synthesis used REML random-effects models where effects were sufficiently comparable; Pearson correlations, rank correlations, ICC/κ/QWK, and other psychometric indices were otherwise retained on metric-appropriate scales. Washback outcomes were synthesized using Hedges' g. RESULTS: Source-level auditing showed that directly comparable Pearson evidence was sparse. Three source-verified L2 studies yielded moderate human-LLM score correspondence (r = 0.66, 95% CI [0.53, 0.75]) with substantial heterogeneity (I 2 = 70.8%). Evidence was strongest in rubric-guided, structurally constrained, and human-supervised contexts, whereas construct representation, fairness, operational reproducibility, and generalizability remained underdeveloped. Across 14 studies, LLM-mediated formative feedback produced a small-to-moderate positive effect on revision quality and short-term writing outcomes (g = 0.42, 95% CI [0.27, 0.57]), although heterogeneity was substantial (I 2 = 69.8%) and concerns about cognitive offloading, learner agency, and reproducibility persisted. DISCUSSION: The findings support a context-dependent interpretation of validity rather than a universal claim about LLM assessment capability. Current evidence supports carefully bounded formative and human-supervised applications but does not justify autonomous high-stakes scoring or broad claims of construct validity, fairness, and generalizability.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Conditional validity in LLM-mediated L2 assessment: an argument-based systematic review and meta-analysis. — 科研速览 Science Skim