科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ JMIR formative research2026-09-22

Effects of Test Length and Scoring Methods on Pass/Fail Classification Accuracy in Single-Choice Examinations: A Simulation Study.

Christiane Pink, Yola Meisel, Arvid Vetter, Philipp Kanzow

一句话结论 · In one sentence

Under the modeled conditions of complete responding, calibrated pass marks, and unchanged response behavior, the investigated scoring methods yielded identical pass/fail classifications. Classification accuracy increased with test length and depended strongly on the distribution of k, with greatest uncertainty near the pass/fail cutoff. These findings apply to the structural effects of the scoring transformations examined here and should not be generalized to behavioral effects of scoring rules in real examinations.

原始摘要(英文原文)· Original abstract
BACKGROUND: Multiple-choice examinations are widely used in dental and medical education. To reduce the effects of random guessing, several scoring methods incorporating penalty scores have been proposed. However, it remains unclear whether such scoring methods improve the accuracy of pass/fail decisions when pass marks are appropriately calibrated. OBJECTIVE: To examine pass/fail classification accuracy under three scoring methods (dichotomous scoring, formula scoring, and right-minus-wrong scoring) in summative examinations using single-choice Type A items with five answer options, and to evaluate the effects of test length and the distribution of the model parameter k (knowledge-based response probability). METHODS: Monte Carlo simulations and analytical derivations were used to evaluate examinations with varying test lengths of up to 300 items. Pass marks were calibrated to the same model-defined cutoff for all scoring methods. Classification accuracy, sensitivity, specificity, and misclassification rates were analyzed across different distributions of k. RESULTS: With every item answered and pass marks calibrated to the same model-defined cutoff, the three scoring methods produced identical pass/fail classifications, as expected from their linear relationship. Classification accuracy increased with test length but was strongly influenced by the distribution of k. Under a uniform distribution, mean accuracy reached 0.90 after 24 items and 0.95 after 94 items, whereas substantially longer tests were required when k was concentrated near the cutoff. These item numbers are model- and distribution-dependent and should not be interpreted as recommended examination lengths. Misclassification was highest near the decision threshold. CONCLUSIONS: Under the modeled conditions of complete responding, calibrated pass marks, and unchanged response behavior, the investigated scoring methods yielded identical pass/fail classifications. Classification accuracy increased with test length and depended strongly on the distribution of k, with greatest uncertainty near the pass/fail cutoff. These findings apply to the structural effects of the scoring transformations examined here and should not be generalized to behavioral effects of scoring rules in real examinations.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Effects of Test Length and Scoring Methods on Pass/Fail Classification Accuracy in Single-Choice Examinations: A Simulation Study. — 科研速览 Science Skim