Ruibing Hou, Mingyue Zhou, Yuwei Gui, Mingshuang Luo, Bingpeng Ma, Hong Chang, Shiguang Shan, Xilin Chen
Faithfully modeling human behavior in dynamic environments is a foundational challenge for embodied intelligence. While conditional motion synthesis has achieved significant advances, egocentric motion generation remains largely underexplored due to the inherent complexity of first-person perception. In this work, we investigate Egocentric Vision-Language (Ego-VL) motion generation. This task requires synthesizing 3D human motion conditioned jointly on first-person visual observations and natural language instructions. We identify a critical optimization challenge in effectively transferring vision-language understanding to motion generation. Directly optimizing vision-language semantic learning and kinematic motion synthesis in an end-to-end manner can lead to optimization interference, limiting the quality of generated motions. To address this challenge, we propose EgoMotion, a two-stage framework for vision-language-guided egocentric motion generation. In the first stage, a vision-language model (VLM) learns motion-aware semantic representations from multimodal inputs through autoregressive motion token prediction. In the second stage, these learned VLM representations serve as expressive conditioning signals for a diffusion-based motion generator. By performing iterative denoising within a continuous latent space, the generator synthesizes physically plausible and temporally coherent trajectories. Extensive evaluations demonstrate that EgoMotion achieves state-of-the-art performance and produces motion sequences that are both semantically grounded and kinematically superior to existing approaches.