Jairo A Navarrete-Ulloa, Valentina Giaconi, Gonzalo Contador
When measurement errors are absent, both n g ¯ and n g ^ are unbiased and any discrepancy between them reflects only sampling variation. When measurement errors are present, n g ¯ acquires a systematic negative bias - consistently underestimating the true learning rate - while n g ^ remains asymptotically unbiased. We further prove that measurement errors induce a spurious negative correlation between pretest scores and ngains, even when prior knowledge and learning capacity are statistically independent.
INTRODUCTION: Nonlinear transformations of pretest and posttest scores are widely used in educational and psychological measurement to estimate group-level change, yet the statistical behavior of estimators derived from such transformations under measurement error remains poorly understood. We examine this problem in the context of normalized gains (ngains), a ratio-based transformation used to estimate group-level "learning rates" in pretest-posttest designs. Two standard estimation methods - the average ngain of the group ( n g ¯ ) and the ngain of the average learner ( n g ^ ) - routinely produce different results. A prior study established a mathematical relationship between this discrepancy and the pretest-ngain correlation, interpreting it as a characterization of the learning process. The pretest-ngain correlation has itself sparked debate: researchers have argued it indicates that ngains favor high-pretest populations, undermining their validity as a measure of student growth.
METHODS: Using Classical Test Theory along with a rencently proposed statistical framework to analize ngains, we show that measurement error is one common cause behind both phenomena.
RESULTS: When measurement errors are absent, both n g ¯ and n g ^ are unbiased and any discrepancy between them reflects only sampling variation. When measurement errors are present, n g ¯ acquires a systematic negative bias - consistently underestimating the true learning rate - while n g ^ remains asymptotically unbiased. We further prove that measurement errors induce a spurious negative correlation between pretest scores and ngains, even when prior knowledge and learning capacity are statistically independent.
DISCUSSION: Such correlations may reflect insufficient instrument reliability rather than any inherent flaw in the transformation. These findings generalize beyond ngains: any nonlinear derived score computed from fallible instruments is susceptible to the same bias structure, and the analytical approach developed here offers a methodological template applicable to other ratio-based metrics in educational and psychological measurement. For applied researchers, we recommend computing both estimators and treating a large discrepancy as a warning sign, reporting instrument reliability alongside ngain estimates, and interpreting pretest-ngain correlations conditionally on reliability.