Jiacheng Qin, Xu Zhang, Dawei Feng, Bo Ding, Yuanzhao Zhai
Benchmarks have established themselves as the standard for evaluating and tracking the progress of Large Language Models. However, the community increasingly faces inconsistent model rankings across nominally similar tasks, suggesting structural misalignments in evaluation instruments. Current paradigms, which rely heavily on aggregate scalar scores, often treat benchmarks as black boxes and fail to diagnose the root causes of these discrepancies. In this work, we propose a structural diagnostic framework grounded in Item Response Theory to evaluate the benchmarks themselves. By mapping items into a latent skill space, we assess benchmark quality along two critical dimensions: calibration, which measures the alignment of difficulty with model capabilities, and efficiency, which evaluates the discriminative power of items. Our analysis of two math benchmarks yields two findings. First, distinct difficulty topologies quantitatively explain why models exhibit conflicting rankings. Second, an observed negative exponential saturation pattern shows that larger item counts can yield diminishing reliability gains. These findings show that quantity does not equal quality, providing a principled roadmap for constructing leaner, more rigorous benchmarks by resolving difficulty mismatches and pruning redundant dimensions.