科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Cureus2026-08-01

A Psychometric Comparison of Faculty-Authored and Large Language Model-Generated Multiple-Choice Questions in Endodontics.

Pradipkumar Damor, Sarita Gill, Wasim Wani, Gauri Arora, Mayank Charan

原始摘要(英文原文)· Original abstract
Background and objective Designing high-quality multiple-choice questions (MCQs) in dental education is highly resource-intensive. While generative artificial intelligence (AI) offers a rapid, scalable solution for automated item generation, there is limited student-tested psychometric evidence comparing uncurated large language model (LLM) outputs with traditional faculty-authored items in advanced dental specialties like endodontics. This study aimed to evaluate and compare the psychometric properties, specifically the difficulty index, discrimination index, distractor effectiveness, and internal consistency reliability, of faculty-authored and uncurated, LLM-generated endodontic MCQs. Materials and methods Sixty single-best-answer MCQs (30 authored by experienced endodontic faculty and 30 generated by the LLM Claude Sonnet 5 (Anthropic, San Francisco, CA) using a standardized single-prompt framework) were distributed evenly across 10 core endodontic domains. The unified 60-item examination was digitally administered to a convenience sample of 100 dental interns in a single, randomized, 60-minute session. Psychometric parameters were computed using classical test theory, and scale reliabilities were evaluated using Cronbach's alpha (α). Continuous variables were compared using independent-samples t-tests. Results The faculty-authored subscale demonstrated good internal consistency (α = 0.851), while the LLM subscale showed acceptable reliability (α = 0.798). Faculty-authored items had a significantly higher mean discrimination capacity (p = 0.00747) and a more balanced difficulty profile, with 27 (90.0%) of the items falling into the moderate difficulty range. Conversely, LLM-generated items were significantly easier (p = 0.000918), with 13 (43.3%) classified as easy, and had a high rate of non-functional distractors (NFDs). While five (16.7%) of the LLM questions had 100% NFDs (where all distractors failed to function), the faculty cohort had no such items. A concurrency analysis revealed that three (10.0%) of the faculty items simultaneously satisfied all ideal psychometric benchmarks, compared to only one (3.3%) of the LLM items. Conclusions Although next-generation LLMs can produce stylistically authentic clinical vignettes that closely mimic human-authored writing, their raw, uncurated outputs exhibit significant psychometric limitations, particularly regarding distractor plausibility and item discrimination. Fully autonomous AI item generation is not yet viable for high-stakes assessments. A hybrid "human-in-the-loop" model, where LLMs are leveraged for rapid preliminary drafting and experienced educators perform targeted distractor refinement, can safeguard assessment validity while substantially reducing the administrative burden on dental faculty.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

A Psychometric Comparison of Faculty-Authored and Large Language Model-Generated Multiple-Choice Questions in Endodontics. — 科研速览 Science Skim