科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Behavior research methods2026-09-24

An LLM canary in the online data coalmine: Bayesian reasoning problems as a capability-gap test for LLM contamination in online samples.

Vera Wilde

原始摘要(英文原文)· Original abstract
The growing use of large language models (LLMs) by participants in online studies threatens the validity of behavioral research data. I propose that Bayesian reasoning problems, which have decades of well-replicated human performance benchmarks, can serve as calibrated detectors of LLM contamination in online samples - an approach I call a capability-gap test. In two preregistered pilots of a Bayesian reasoning training tool (N = 148), participants' accuracy on positive predictive value (PPV) calculation problems reached roughly three times established human performance ceilings. Specifically, 57% of Pilot 2 participants achieved perfect 5/5 scores, against a meta-analytic ceiling of approximately 24% for a single problem presented in natural frequency format, a level at which perfect scores are effectively unattainable. Including two outcome measures - accuracy (correct numerical answer) and Bayesian algorithm use (evidence of the reasoning process) - allows researchers to distinguish LLM contamination from authentic learning effects: accuracy detects contamination, while algorithm use preserves treatment effect signals. This is itself a signal detection problem, structurally analogous to the mass screenings for low-prevalence problems around which the Bayesian reasoning literature was developed. Bayesian reasoning problems offer four advantages as contamination detectors: well-established human performance ceilings from meta-analyses, a large human-LLM performance gap, known reference distributions enabling nuanced assessment, and ease of embedding in existing surveys. I provide practical recommendations for using these problems as data quality diagnostics in online research.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

An LLM canary in the online data coalmine: Bayesian reasoning problems as a capability-gap test for LLM contamination in online samples. — 科研速览 Science Skim