Walter P Vispoel, Hyeryung Lee, Tingting Chen
Researchers have long known that scores from Likert-style self-report measures of psychological traits are affected by both item wording (negative and positive) and multiple sources of measurement error (specific-factor, transient, and random-response), but the magnitude of such effects is rarely separated and individually estimated within research studies. When these effects are not distinguished, as is the case when reporting conventional reliability estimates (e.g., alpha, omega, split-half, test-retest), they are confounded with trait variance and/or each other, typically resulting in overestimation of score accuracy. We demonstrate how to quantify and separate trait, wording, and multiple measurement error effects within extended bifactor models using data from a large sample of respondents (n = 1796) who completed the Physical Appearance subscale from the Self-Description Questionnaire-III on two occasions. Results revealed that models including item wording effects for both single and multiple occasion designs provided noticeably superior fits to the data in comparison to models excluding such effects, with item wording accounting for 8.5% to 8.8% of total score variance across designs. We further demonstrate how to extend analyses to the individual item level to identify items that are most and least affected by wording and measurement error effects and provide code in R to enable readers to apply these techniques to their own data.