Shawn Hemelstrand, Tomohiro Inoue
It is common in the language sciences to dichotomize continuous data in order to fit models to data. However, several statisticians and methodologists have warned against this practice for years. Many in the language sciences seem unaware of this problem. Because of the lack of modern, robust, and open data simulations related to this issue in the language science literature, this article provides an empirical investigation of this practice. Across three different simulations, our analysis shows that dichotomization almost universally increases the standard errors, and consequently leads to inaccuracy of tests of statistical significance. Furthermore, effect sizes like R 2 are often diminished by the reduction of available information in the data. We conclude by providing suggestions and considerations for future empirical studies.