科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Frontiers in psychology2026-01-01

Paper- vs. computer-based assessment in high-stakes testing: a common-person investigation of mode comparability.

Kübra Atalay Kabasakal, Tuba Gündüz

原始摘要(英文原文)· Original abstract
This study examined the mode comparability of paper-based (YDS) and computer-based (e-YDS) forms of a high-stakes foreign language examination administered in Türkiye. Rather than treating the analysis as a routine test equating, we framed it as a test of whether a change in administration mode leaves the measured construct invariant and the resulting scores comparable. Using a common-person design and the Rasch model, six form pairs were placed on a common scale through both separate calibration with mean/sigma transformation and concurrent calibration. Model-data fit and dimensionality (via tetrachoric inter-item correlations and parallel analysis) were evaluated for every form, and the linking was examined with and without the exclusion of inconsistent examinees as a sensitivity analysis. Scaling constants, the linking RMSE, and the paired mean difference (Δ) were computed, and cross-form test information functions were compared. Across all pairs the data were essentially unidimensional and fit the Rasch model adequately, with a comparably dominant single dimension for paper-based and computer-based forms. Although the computer-based forms did not display a stronger secondary dimension than the paper-based forms, the design-with different items and administration dates-cannot isolate a pure mode effect. Linking constants and the estimated mode difference were nearly identical with and without examinee exclusion (mode differences ≤ ~0.17 logits), and the separate and concurrent methods agreed almost perfectly at the item-parameter level (r = 1.00). Computer-based forms yielded slightly higher ability estimates, but the common-examinee distributions remained stable. Taken together, the results provide evidence of practical mode comparability under the present operational conditions, while we are explicit about the confounds and assumptions that bound this conclusion.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Paper- vs. computer-based assessment in high-stakes testing: a common-person investigation of mode comparability. — 科研速览 Science Skim