Sergio Escorial
Intelligence test scores underpin high-stakes educational, clinical, and occupational decisions, which makes the metric quality of these instruments a central concern for professional practice. This study aimed to evaluate the psychometric and documentary quality of intelligence tests reviewed by the Spanish National Test Commission, compare this quality with that of tests measuring other psychological constructs, and examine the influence of publisher and evaluation cohort on quality. To this end, 37 intelligence tests and 78 tests of other constructs, reviewed across 13 editions (2010-2026) of the project, were coded using the Revised Test Review Questionnaire (CET-R), which comprises 14 quality characteristics grouped into four dimensions: General, Validity, Reliability, and Norms. Intelligence tests did not report more information than tests of other constructs, showing similar omission rates. However, when information was reported, intelligence tests obtained higher scores on the General, Validity, and Norms dimensions, with no differences in Reliability. Publisher accounted for a substantial share of the variance in quality, whereas the evaluation cohort showed no significant influence on any psychometric characteristic once corrected for multiple comparisons (FDR). Item bias analyses and item-response-theory-based reliability remained the most consistently underreported forms of evidence across the manuals of all tests analyzed. These findings underscore the need to uphold high, transparent standards in the documentation of intelligence tests and provide empirical criteria to support the responsible selection of instruments by practitioners.