Xiting Wang, Liming Jiang, José Hernández‐Orallo, David Stillwell, Shiqiang Chen, Luning Sun, Fang Luo, Xing Xie
Rigorous evaluation of general-purpose AI systems such as large language models should allow for deepened understanding of their capabilities and effective mitigation of their risks. The current evaluation paradigm, mostly reliant on benchmarks aggregating scores on one or more tasks, lacks the scientific machinery for predicting performance on unforeseen tasks and explaining the variability of results. Moreover, existing benchmarks raise growing concerns about their reliability and validity. To tackle these challenges, we vindicate psychometrics, the science of psychological measurement, as a methodology for identifying and measuring constructs that underlie AI performance across multiple tasks. To raise awareness, we first identify the key advantages of adapting psychometric principles to AI evaluation through concrete examples; second, we distinguish sound applications of psychometric techniques from oversimplified ones and warn against common pitfalls; and third, to encourage general use, we introduce a systematic psychometric framework and an operational evaluation pipeline, which provide practical implementation guidance. In the end, we discuss underexplored avenues and societal implications that open new research directions for the use of psychometrics in broader AI research.