Vera Wilde
The growing use of large language models (LLMs) by participants in online studies threatens the validity of behavioral research data. I propose that Bayesian reasoning problems, which have decades of well-replicated human performance benchmarks, can serve as calibrated detectors of LLM contamination in online samples - an approach I call a capability-gap test. In two preregistered pilots of a Bayesian reasoning training tool (N = 148), participants' accuracy on positive predictive value (PPV) calculation problems reached roughly three times established human performance ceilings. Specifically, 57% of Pilot 2 participants achieved perfect 5/5 scores, against a meta-analytic ceiling of approximately 24% for a single problem presented in natural frequency format, a level at which perfect scores are effectively unattainable. Including two outcome measures - accuracy (correct numerical answer) and Bayesian algorithm use (evidence of the reasoning process) - allows researchers to distinguish LLM contamination from authentic learning effects: accuracy detects contamination, while algorithm use preserves treatment effect signals. This is itself a signal detection problem, structurally analogous to the mass screenings for low-prevalence problems around which the Bayesian reasoning literature was developed. Bayesian reasoning problems offer four advantages as contamination detectors: well-established human performance ceilings from meta-analyses, a large human-LLM performance gap, known reference distributions enabling nuanced assessment, and ease of embedding in existing surveys. I provide practical recommendations for using these problems as data quality diagnostics in online research.