科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Research Methods in Applied Linguistics2025-10-15· Computer science

Stop splitting hairs: The problems with dichotomizing continuous data in language research

Shawn Hemelstrand, Tomohiro Inoue

原始摘要(英文原文)· Original abstract
It is common in the language sciences to dichotomize continuous data in order to fit models to data. However, several statisticians and methodologists have warned against this practice for years. Many in the language sciences seem unaware of this problem. Because of the lack of modern, robust, and open data simulations related to this issue in the language science literature, this article provides an empirical investigation of this practice. Across three different simulations, our analysis shows that dichotomization almost universally increases the standard errors, and consequently leads to inaccuracy of tests of statistical significance. Furthermore, effect sizes like R 2 are often diminished by the reduction of available information in the data. We conclude by providing suggestions and considerations for future empirical studies.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Stop splitting hairs: The problems with dichotomizing continuous data in language research — 科研速览 Science Skim