Kübra Atalay Kabasakal, Tuba Gündüz
This study examined the mode comparability of paper-based (YDS) and computer-based (e-YDS) forms of a high-stakes foreign language examination administered in Türkiye. Rather than treating the analysis as a routine test equating, we framed it as a test of whether a change in administration mode leaves the measured construct invariant and the resulting scores comparable. Using a common-person design and the Rasch model, six form pairs were placed on a common scale through both separate calibration with mean/sigma transformation and concurrent calibration. Model-data fit and dimensionality (via tetrachoric inter-item correlations and parallel analysis) were evaluated for every form, and the linking was examined with and without the exclusion of inconsistent examinees as a sensitivity analysis. Scaling constants, the linking RMSE, and the paired mean difference (Δ) were computed, and cross-form test information functions were compared. Across all pairs the data were essentially unidimensional and fit the Rasch model adequately, with a comparably dominant single dimension for paper-based and computer-based forms. Although the computer-based forms did not display a stronger secondary dimension than the paper-based forms, the design-with different items and administration dates-cannot isolate a pure mode effect. Linking constants and the estimated mode difference were nearly identical with and without examinee exclusion (mode differences ≤ ~0.17 logits), and the separate and concurrent methods agreed almost perfectly at the item-parameter level (r = 1.00). Computer-based forms yielded slightly higher ability estimates, but the common-examinee distributions remained stable. Taken together, the results provide evidence of practical mode comparability under the present operational conditions, while we are explicit about the confounds and assumptions that bound this conclusion.