科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ JMIR medical education2026-09-09

AI-Assisted Angoff Standard Setting for Multiple-Choice Examinations in Medical Education: Comparative Study.

Salah Eldin Kassab, Mariam Shadan, Arina Ziganshina, Shifan Khanday, Hossam Hamdy

一句话结论 · In one sentence

In this single-institution study, the evaluated LLMs generated modified Angoff estimates that were broadly comparable to those of faculty judges. Generalizability theory analyses demonstrated higher G and Φ coefficients, lower rater-related variance, and lower RMSE for the evaluated LLM outputs under standardized prompting conditions, indicating greater consistency in the generated estimates. Application of the resulting cutoff scores produced pass and fail rates similar to those derived from faculty judges. These findings suggest the future application of LLMs as decision assistance tools for modified Angoff standard setting while maintaining expert human oversight.

原始摘要(英文原文)· Original abstract
BACKGROUND: Standard setting is essential for a defensible assessment in medical education. The modified Angoff method requires several expert judges, and determining the minimally competent candidate is cognitively challenging. However, empirical evidence on the role of AI in standard setting is unclear. OBJECTIVE: This study aimed to examine the role of large language models (LLMs) in the modified Angoff standard setting method for multiple-choice questions in a basic medical science examination compared with faculty judges. METHODS: This study was conducted in year 2 of the MD program in the United Arab Emirates. Ten faculty judges and 5 LLMs (GPT-5.2, Grok, DeepSeek, MedGemma, and Claude 4.5 Sonnet) determined the Angoff cutoff scores for a summative examination (120 multiple-choice questions). Standardized prompts were used for the LLMs to mimic the same information provided for faculty judges. Generalizability (G) theory analyses were performed using a fully crossed item × rater design to estimate variance components, G and Φ coefficients, decision study, and root mean squared error (RMSE) of the Angoff cutoff scores. RESULTS: Faculty-generated Angoff estimates (mean 69.13, SD 10.92) were comparable to LLM-generated estimates (mean 68.94, SD 12.39). A 2-tailed paired-sample t test revealed no statistically significant difference between the 2 groups (95% CI -2.21 to 2.91; t119=0.27; P=.79). Generalizability theory analysis demonstrated moderate reliability for the faculty panel (G coefficient=0.738; Φ coefficient=0.692). Despite comprising only 5 LLMs, the LLM panel demonstrated higher reliability (G coefficient=0.823; Φ coefficient=0.815), lower RMSE (1.28 vs 2.33), higher item-related variance (46.85% vs 18.36%), and lower rater-related variance (2.64% vs 16.40%) than the faculty panel. Pass rates were similar using LLM- and faculty-derived cutoff scores (52/73, 71.2%). LLMs differed in their minimally competent candidate conceptualization and approaches to determining item-level percentages of correct responses. Furthermore, the correlations between Angoff estimates and item-related P values were larger in LLMs than in faculty judges (r=0.552 vs 0.437). CONCLUSIONS: In this single-institution study, the evaluated LLMs generated modified Angoff estimates that were broadly comparable to those of faculty judges. Generalizability theory analyses demonstrated higher G and Φ coefficients, lower rater-related variance, and lower RMSE for the evaluated LLM outputs under standardized prompting conditions, indicating greater consistency in the generated estimates. Application of the resulting cutoff scores produced pass and fail rates similar to those derived from faculty judges. These findings suggest the future application of LLMs as decision assistance tools for modified Angoff standard setting while maintaining expert human oversight.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

AI-Assisted Angoff Standard Setting for Multiple-Choice Examinations in Medical Education: Comparative Study. — 科研速览 Science Skim