Daniel Gianola, Olga Ravagnolo, Chris-Carolin Schoen
It is common in quantitative genetics to submit phenotypes to pre-processing prior to data analysis. The objective is to make computations simpler. In areas where field data is observational, e.g., studies of complex traits, records are often pre-adjusted for nuisance factors such as location, demographical structure and macroenvironments, that may mask expression of genetic values. The nuisance effects may be estimated by ordinary or generalized least-squares. We studied consequences of ignoring pre-correction of nuisance (fixed) effects using a mixed linear model framework supplemented by analyses of real and synthetic plant and animal breeding data sets. Pre-correction induces changes in model rank and sometimes distorts inferences. We also examined prediction of random effects with pre-corrected data in situations with dense or sparse information on the levels of fixed effects. Variance component analyses with ANOVA, minimum norm quadratic unbiased estimation), maximum likelihood and Bayesian methods were also dealt with. A Braford cattle data set with birth weight of calves as response variable illustrated effects on prediction of phenotypes. Our main conclusions are: 1) the impact of ignoring pre-correction depends on model complexity relative to sample size. 2) Pre-processing must be used with caution in inference as it may destroy data structure of the data and obscure the understanding of a problem under investigation.