科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Communications of the ACM2026-04-14· Computer science

Evaluating General-Purpose AI with Psychometrics

Xiting Wang, Liming Jiang, José Hernández‐Orallo, David Stillwell, Shiqiang Chen, Luning Sun, Fang Luo, Xing Xie

原始摘要(英文原文)· Original abstract
Rigorous evaluation of general-purpose AI systems such as large language models should allow for deepened understanding of their capabilities and effective mitigation of their risks. The current evaluation paradigm, mostly reliant on benchmarks aggregating scores on one or more tasks, lacks the scientific machinery for predicting performance on unforeseen tasks and explaining the variability of results. Moreover, existing benchmarks raise growing concerns about their reliability and validity. To tackle these challenges, we vindicate psychometrics, the science of psychological measurement, as a methodology for identifying and measuring constructs that underlie AI performance across multiple tasks. To raise awareness, we first identify the key advantages of adapting psychometric principles to AI evaluation through concrete examples; second, we distinguish sound applications of psychometric techniques from oversimplified ones and warn against common pitfalls; and third, to encourage general use, we introduce a systematic psychometric framework and an operational evaluation pipeline, which provide practical implementation guidance. In the end, we discuss underexplored avenues and societal implications that open new research directions for the use of psychometrics in broader AI research.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Evaluating General-Purpose AI with Psychometrics — 科研速览 Science Skim