Nader Alnomasy, Habib Alrashedi, Maha Dardouri, Soha Kamel Mosbah Mahmoud, Petelyne Pangket, Sang Bin You, Aluem Tark, Deborah Becker, Richard Balacuit Maestrado, Romeo Mostoles, Rayhanah R Almutairi, Ebtsam Abouhashish, Sudharani B Banappagoudar, Asim A Alreshidi, Waleed Meajib Alshammari, Faisal Fahaad Alrshidi, Razan Alsayed, Sharifah Alsayed, Jiyoun Song
Background/Objectives: High-quality multiple-choice questions (MCQs) are essential for valid assessment in graduate-level nursing programs, yet developing such items is resource-intensive for educators. Large language models such as ChatGPT may generate educational assessment content; however, evidence regarding their effectiveness in producing graduate-level nursing examination items remains limited. This study compared the perceived quality of ChatGPT-generated and educator-authored MCQs for graduate-level nursing examinations and examined exploratory associations between assessor characteristics and quality ratings. Methods: A comparative cross-sectional study was conducted among 27 international nurse educators who independently evaluated 50 MCQs (25 ChatGPT-generated and 25 educator-authored) using an expert-reviewed rubric assessing eight quality domains, with the primary outcome being the between-source difference in perceived quality ratings. Participants were blinded to question source. Results: ChatGPT-generated MCQs received nominally higher ratings than educator-authored items across most domains, but effect sizes were uniformly small (Cohen's d = 0.11-0.33). No significant differences were observed in perceived difficulty or distractor plausibility; after Bonferroni correction for multiple domain comparisons, differences in cognitive level, scenario relevance, and instructional alignment remained significant. Conclusions: ChatGPT-generated MCQs achieved perceived-quality ratings broadly comparable to educator-authored items when supported by structured prompting and expert refinement, although expert review remains necessary before classroom use, and these results should be regarded as preliminary. These findings support AI-assisted item development as a complementary strategy pending replication in larger, more rigorous studies.