科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Journal of voice : official journal of the Voice Foundation2026-08-25

Online, Crowdsourced Sampling (OCS) Platforms for Large-Scale Data Collection in Voice Disorders Research: Prices, Pitfalls, and Perks.

Christopher S Apfelbach, Lady Catherine Cantor-Cutiva, Eric J Hunter

一句话结论 · In one sentence

Content-based screenings, particularly analysis of responses to open-ended questions, meaningfully augmented automated screenings. Low-quality respondents systematically over-reported health- and voice-related disability, a severity bias with plausible economic and behavioral explanations. The emergence of AI-generated responses represents an evolving and distinct challenge that content-based screening may be increasingly insufficient to address. We provide recommendations for collecting high-quality data on OCS platforms to minimize barriers to entry for future clinical voice research.

原始摘要(英文原文)· Original abstract
OBJECTIVE: Online, crowdsourced sampling (OCS) platforms such as Amazon's Mechanical Turk and CloudResearch Connect enable faster, cheaper, and larger-scale data collection than in-person sampling methods. However, OCS platforms are often criticized due to data quality concerns. This study examines a battery of screening tools to determine which best predicts data quality and to quantify the influence of low-quality data on the statistical properties of the samples. Finally, we issue recommendations to researchers interested in OCS-based clinical voice research that may reduce costly missteps when fielding their first OCS studies. METHODS: To evaluate the suitability of OCS for clinical voice research, survey-based measures of vocal function, vocal fatigue, personality, communicative quality of life, and other factors were collected on Qualtrics from (1) a convenience sample of undergraduate students (n = 47), (2) Mechanical Turk users (n = 495), and (3) Connect users (n = 99) over 6 months. Low-quality responses were eliminated using six screening tools in four nested stages. Logistic regression was used to identify variables that strongly predicted data quality. Finally, patient-reported outcome measure (PROM) responses were compared between the high- and low-quality groups. RESULTS: Depending on the stage, data quality screenings flagged between 5.00% (n = 32) and 13.6% (n = 87) of responses as low-quality. Hispanic/Latino ethnicity and long response times most strongly predicted low-quality responses, which uniformly exhibited more severe ratings of health- and voice-related disability than did high-quality responses. CONCLUSIONS: Content-based screenings, particularly analysis of responses to open-ended questions, meaningfully augmented automated screenings. Low-quality respondents systematically over-reported health- and voice-related disability, a severity bias with plausible economic and behavioral explanations. The emergence of AI-generated responses represents an evolving and distinct challenge that content-based screening may be increasingly insufficient to address. We provide recommendations for collecting high-quality data on OCS platforms to minimize barriers to entry for future clinical voice research.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Online, Crowdsourced Sampling (OCS) Platforms for Large-Scale Data Collection in Voice Disorders Research: Prices, Pitfalls, and Perks. — 科研速览 Science Skim