Abdul Hannan Zulkarnain, Attila Gere
Immersive virtual reality (VR) is increasingly used in food sensory research, but participant burden and exclusions complicate sample-size planning. We benchmark analysed sample sizes in immersive head-mounted display (HMD) VR sensory evaluation using a Scopus-only structured search (2021–2025). The search yielded 86 records; after limiting to research articles, 45 records were screened in two stages (title/abstract, then full text), 45 full texts were assessed, and 15 independent experiments met inclusion criteria. Studies were labelled Positive if they reported at least one statistically significant ( p < 0.05) VR effect on a main food response outcome (any-positive rule), and Negative otherwise. Across eligible studies ( n = 15), analysed sample sizes ranged from 15 to 279 (median 63). Eight studies were Positive (8/15), giving an eligible-only probability of 0.5333 (Wilson 95% CI 0.3012–0.7519). With a transparent baseline Beta(1,1) prior, the posterior mean was 0.5294 (95% CrI 0.2988–0.7535). A participant-weighted estimate based on analysed N was 0.4864. A hierarchical Beta–Binomial model ( μ ∼ Beta(1,1); kappa ∼ Exponential(1)) estimated μ =0.5117 (95% CrI 0.1177–0.8982) and kappa = 1.5456 (95% CrI 0.1588–4.6401). We define the minimum defensible analysed N as 49.25 (25th percentile of analysed N among Positive studies), used as a pragmatic lower-quartile benchmark balancing feasibility and defensibility while avoiding reliance on extreme minima. This supports planning for ≈50 analysed participants for typical within-subject VR sensory designs in this evidence base; between-subject designs should be planned per group and, in this evidence base, used total analysed N ≥ 100.