Chaoqun Li, Longfei Wang, Huiying Xu, Xiaolei Zhang, Zhenglong Wan, Xinzhong Zhu
Embodied picking for e-commerce fulfillment remains vulnerable to dense clutter, reflective packaging, and background variation, which can undermine the effectiveness and robustness of visuomotor policies learned from demonstrations. A key limitation is the absence of explicit object grounding, causing policies to exploit spurious contextual cues rather than task-relevant visual evidence. To address this issue, we propose the Foveated Diffusion Policy (FDP), which integrates object-centric visual grounding into diffusion-based action generation. FDP adopts a panorama-fovea dual-stream visual encoder: a low-resolution panorama stream captures global scene context, while a detection-guided fovea stream uses ROI Align to extract high-resolution object tokens. At each denoising step, these tokens are incorporated into the diffusion denoiser through object-guided cross-attention, thereby grounding action generation in target-specific visual evidence instead of undifferentiated global representations. We evaluate FDP on ten benchmark datasets spanning six simulated manipulation tasks with both proficient-human and multi-human demonstrations. FDP achieves average success rates of 81.0% and 76.8%, respectively, outperforming diffusion-based baselines and delivering particularly strong performance on object-centric tasks, including Push-T (86.7 ± 2.5%) and Lift (100.0 ± 0.0%). Ablation studies on a diagnostic subset show that foveated conditioning improves the average success rate by 11.3 percentage points over the non-foveated variant. FDP also maintains an 84.0% success rate under severe localization perturbations (σ = 10 pixels), demonstrating resilience to detection noise. Furthermore, accelerated DDIM sampling reduces policy-inference latency to approximately 55 ms with 10 denoising steps, supporting the feasibility of closed-loop control. Collectively, these results demonstrate that explicit object grounding is a promising and computationally practical approach to improving the robustness of visuomotor policies in warehouse-relevant simulated manipulation scenarios.