科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Frontiers in robotics and AI2026-01-01

Object-grounded embodied picking for e-commerce warehouse fulfillment: a foveated diffusion policy for operational robustness.

Chaoqun Li, Longfei Wang, Huiying Xu, Xiaolei Zhang, Zhenglong Wan, Xinzhong Zhu

原始摘要(英文原文)· Original abstract
Embodied picking for e-commerce fulfillment remains vulnerable to dense clutter, reflective packaging, and background variation, which can undermine the effectiveness and robustness of visuomotor policies learned from demonstrations. A key limitation is the absence of explicit object grounding, causing policies to exploit spurious contextual cues rather than task-relevant visual evidence. To address this issue, we propose the Foveated Diffusion Policy (FDP), which integrates object-centric visual grounding into diffusion-based action generation. FDP adopts a panorama-fovea dual-stream visual encoder: a low-resolution panorama stream captures global scene context, while a detection-guided fovea stream uses ROI Align to extract high-resolution object tokens. At each denoising step, these tokens are incorporated into the diffusion denoiser through object-guided cross-attention, thereby grounding action generation in target-specific visual evidence instead of undifferentiated global representations. We evaluate FDP on ten benchmark datasets spanning six simulated manipulation tasks with both proficient-human and multi-human demonstrations. FDP achieves average success rates of 81.0% and 76.8%, respectively, outperforming diffusion-based baselines and delivering particularly strong performance on object-centric tasks, including Push-T (86.7 ± 2.5%) and Lift (100.0 ± 0.0%). Ablation studies on a diagnostic subset show that foveated conditioning improves the average success rate by 11.3 percentage points over the non-foveated variant. FDP also maintains an 84.0% success rate under severe localization perturbations (σ = 10 pixels), demonstrating resilience to detection noise. Furthermore, accelerated DDIM sampling reduces policy-inference latency to approximately 55 ms with 10 denoising steps, supporting the feasibility of closed-loop control. Collectively, these results demonstrate that explicit object grounding is a promising and computationally practical approach to improving the robustness of visuomotor policies in warehouse-relevant simulated manipulation scenarios.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Object-grounded embodied picking for e-commerce warehouse fulfillment: a foveated diffusion policy for operational robustness. — 科研速览 Science Skim