科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Lecture notes in computer science2026-08-02· Computer science

Diagnosing LLM Benchmark: A Psychometric Analysis of Difficulty and Discrimination

Jiacheng Qin, Xu Zhang, Dawei Feng, Bo Ding, Yuanzhao Zhai

原始摘要(英文原文)· Original abstract
Benchmarks have established themselves as the standard for evaluating and tracking the progress of Large Language Models. However, the community increasingly faces inconsistent model rankings across nominally similar tasks, suggesting structural misalignments in evaluation instruments. Current paradigms, which rely heavily on aggregate scalar scores, often treat benchmarks as black boxes and fail to diagnose the root causes of these discrepancies. In this work, we propose a structural diagnostic framework grounded in Item Response Theory to evaluate the benchmarks themselves. By mapping items into a latent skill space, we assess benchmark quality along two critical dimensions: calibration, which measures the alignment of difficulty with model capabilities, and efficiency, which evaluates the discriminative power of items. Our analysis of two math benchmarks yields two findings. First, distinct difficulty topologies quantitatively explain why models exhibit conflicting rankings. Second, an observed negative exponential saturation pattern shows that larger item counts can yield diminishing reliability gains. These findings show that quantity does not equal quality, providing a principled roadmap for constructing leaner, more rigorous benchmarks by resolving difficulty mismatches and pruning redundant dimensions.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Diagnosing LLM Benchmark: A Psychometric Analysis of Difficulty and Discrimination — 科研速览 Science Skim