科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ International Journal of Computational Intelligence Systems2026-07-31· Computer science

Systematic Survey of Large Language Model Evaluation: Methods, Benchmarks, and Comprehensive Taxonomic Framework

Nadir iBRAHIMOĞLU, Mustafa Umut Demi̇rezen, Furkan Yıldız

原始摘要(英文原文)· Original abstract
Large Language Models (LLMs) have rapidly reshaped research and practice in natural language processing, but methods for evaluating these systems have developed more slowly. In this survey we review work on LLM evaluation published between 2020 and 2025, drawing on 182 studies identified through a PRISMA 2020–guided search of five major sources (ACL Anthology, IEEE Xplore, ACM Digital Library, arXiv, and Google Scholar) and subsequent quantitative and qualitative analysis. On the basis of this literature, we propose a three-dimensional framework that groups evaluation methods by methodology (intrinsic, extrinsic, and human-based), by scope (task-specific, general-purpose, and domain-specific), and by the Evaluation Framework Approach employed (automated metrics, benchmark suites, and interactive/hybrid assessment). We use this framework to characterize current practice and to highlight three recurring difficulties: bias introduced by evaluation design choices, rapid saturation of popular benchmarks, and weak correspondence between standard metrics and performance in real use cases. We then examine how widely used evaluation frameworks such as HELM, LM-Evaluation Harness, and TruLens partially address, but in some respects also reproduce, these limitations, particularly in multilingual, culturally diverse, and deployed settings. The survey concludes with practical recommendations for designing and reporting LLM evaluations, guidance for practitioners selecting or building evaluation pipelines, and a research agenda that emphasizes closer alignment between evaluation protocols and the conditions under which LLMs are actually used.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Systematic Survey of Large Language Model Evaluation: Methods, Benchmarks, and Comprehensive Taxonomic Framework — 科研速览 Science Skim