Kelechi Nnaemeka Chukwudi, Brian Hallahan, Colm McDonald
Modified Angoff scores generally do predict group-level "borderline-pass" students' performance adequately, however their accuracy varies by question type, content area and psychometric property. Further standard-setter training, and inclusion of more diverse specialist input may improve standard-setting accuracy across less familiar domains.
OBJECTIVES: Multiple-choice questions (MCQs) are widely used in medical education, with the Angoff standard-setting method frequently used to determine competence. However, minimal research has determined the accuracy of item-level Angoff scores in predicting the actual performance of borderline-pass students.
METHODS: This retrospective study analysed five years (2019-2023) of psychiatry MCQ data (350 questions) from 191 fourth-year borderline-pass undergraduate medical students (scoring 45-55% in their psychiatry module MCQ) at the University of Galway. Angoff-derived cut scores were compared with actual student scores across clinical and knowledge-based questions, diagnostic categories, and content domains.
RESULTS: Angoff (mean = 53.0%) and actual student scores (mean = 54.2%) showed minimal overall difference and were moderately correlated (r = 0.57, p < 0.001). However, questions where clinical scenarios were presented were underestimated (-6.2%), and knowledge-based questions overestimated (+2.9%) by standard-setters. Question difficulty was not a significant predictor of score differences. Lower discrimination scores predicted standard setter marks being higher than student marks for knowledge-based (B = -31.96, β = -0.19, t = 2.61, p = 0.01) questions.
CONCLUSION: Modified Angoff scores generally do predict group-level "borderline-pass" students' performance adequately, however their accuracy varies by question type, content area and psychometric property. Further standard-setter training, and inclusion of more diverse specialist input may improve standard-setting accuracy across less familiar domains.