Artyom Zinchenko, Lingyue Chen, Zhuanghua Shi, Thomas Geyer, Loes van Dam
Repeated scenes facilitate visual search, but standard contextual-cueing paradigms collapse multiple processing stages into a single reaction time. Using immersive virtual reality and continuous head tracking (N = 100), we decomposed memory-guided search into Global Orienting Time (GOT), Local Alignment Time (LAT), and Target Identification Time (TIT). After a learning phase, a 180° scene rotation reversed each repeated target's egocentric position while preserving its local distractor context; half the participants physically rotated their chairs, providing proprioceptive cues to a viewpoint change. For targets initially outside the field of view (FOV), the contextual effect on GOT changed across phases: no benefit during learning and a small numerical cost for repeated contexts after rotation. In contrast, LAT was reliably facilitated by repeated contexts after rotation, whereas TIT was only numerically faster for repeated displays in both phases. For targets initially inside the FOV, which required no global orienting, the LAT benefit observed during learning reversed into a cost after rotation, whereas TIT remained faster for repeated displays throughout; the net RT benefit was abolished rather than reversed. Thus, the same manipulation produced opposing effects on earlier and later stages of a single search episode, partially cancelling in aggregate RTs. Physical rotation did not modulate any contextual-cueing effect, providing no evidence that self-motion cues updated the stored representation. These findings point to a spatial architecture of contextual memory in which viewpoint-dependent and display-centered contributions are expressed at distinct stages of search and are obscured in aggregate RTs.