María J Blanca, Rafael Alarcón, Roser Bono, Jaume Arnau, F Javier García-Castro, Guillermo Vallejo
With balanced designs, the results show that for the group effect the F-statistic and B-F are valid choices under assumption violations. For time and interaction effects, under non-sphericity, F-GG and F-HF maintain Type I error close to 5% with moderate violation of normality; both of them remain robust with ε ^ ≥ 0.70 for larger violations of normality. With unbalanced designs, the behavior of all these statistics depends on the variables manipulated.
INTRODUCTION: Data analysis of split-plot designs requires normality, homogeneity of variance, and multisample sphericity. Adjusted F-tests and the bootstrap-F procedure have been proposed as alternatives to the F-statistic when these assumptions are violated. However, simulation studies have not identified the conditions under which these tests are robust.
METHOD: To analyze the Type I error of the F-statistic, Greenhouse-Geisser (F-GG) and Huynh-Feldt (F-HF) adjustments, and bootstrap-F (B-F), we performed a simulation study with split-plot designs, manipulating the distribution shape, epsilon values, sample size, coefficient of sample size variation, degree of heterogeneity, and pairing of variance with group sample size.
RESULTS: With balanced designs, the results show that for the group effect the F-statistic and B-F are valid choices under assumption violations. For time and interaction effects, under non-sphericity, F-GG and F-HF maintain Type I error close to 5% with moderate violation of normality; both of them remain robust with ε ^ ≥ 0.70 for larger violations of normality. With unbalanced designs, the behavior of all these statistics depends on the variables manipulated.
DISCUSSION: The paper identifies the conditions under which each procedure can be used. Overall, B-F was the most robust procedure in the majority of conditions. It can be used under moderate violations of normality, homogeneity, and sphericity with N > 20. Under more severe violations of these assumptions, larger sample size, N > 80, is needed to achieve robustness.