科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Frontiers in neurorobotics2026-01-01

GSP-3D: generalizable 3D Gaussian Splatting with diffusion policy and world model for long-horizon bimanual manipulation.

Xukun Liu, Junhua Huang, Shenggang Wei, Liguo Hou, Junmou Qiao, Aifeng Liu, Hongsheng Huang, Fengjuan Xie

一句话结论 · In one sentence

We evaluate GSP-3D on the RoboTwin 2.0 benchmark across a range of bimanual manipulation tasks under both clean and domain-randomized conditions, including variations in lighting, backgrounds, and tabletop distractors. Experimental results show that GSP-3D consistently outperforms existing baseline methods, achieving higher success rates while maintaining minimal computational overhead.

原始摘要(英文原文)· Original abstract
INTRODUCTION: Learning robust visuomotor policies for bimanual manipulation remains challenging due to the stringent requirements for precise coordination between arms and the ability to generalize across diverse environmental conditions. Existing diffusion-based policies often suffer from temporally inconsistent action generation, while their reliance on sparse point cloud representations limits structural completeness and fails to capture future scene dynamics, hindering performance in long-horizon tasks. METHODS: To address these limitations, we introduce GSP-3D, a unified framework that integrates generalizable 3D Gaussian Splatting (3DGS) with diffusion-based policy learning. GSP-3D comprises three key components: (1) a Generalizable Gaussian Regressor (GGR) that predicts 3D Gaussian primitives from a single RGB-D frame in real time; (2) a 3DGS-conditioned diffusion policy that aggregates Gaussians into compact latent representations, replacing point clouds with explicit geometric primitives; and (3) a transformer-based world model that forecasts future Gaussian sets and leverages prediction error as an auxiliary loss to enforce temporal consistency across actions. RESULTS: We evaluate GSP-3D on the RoboTwin 2.0 benchmark across a range of bimanual manipulation tasks under both clean and domain-randomized conditions, including variations in lighting, backgrounds, and tabletop distractors. Experimental results show that GSP-3D consistently outperforms existing baseline methods, achieving higher success rates while maintaining minimal computational overhead. DISCUSSION: These findings demonstrate that integrating explicit 3D Gaussian representations with diffusion policies offers an efficient and robust solution for long-horizon, temporally coherent bimanual manipulation, effectively addressing the generalization and consistency challenges that limit current approaches.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

GSP-3D: generalizable 3D Gaussian Splatting with diffusion policy and world model for long-horizon bimanual manipulation. — 科研速览 Science Skim