Haibo Yang, Yang Chen, Yingwei Pan, Zhineng Chen, Ting Yao, Tao Mei
Autoregressive models via next token prediction have emerged as a promising way towards unified models across multiple modalities for general artificial intelligence. Nevertheless, exploiting this idea to achieve single image to 3D generation remains a challenging problem. In this paper, we introduce a new Autoregressive 3D diffusion model (dubbed as AR3D), that novelly frames autoregressive modeling on multi-view images of 3D content as "next symmetric view prediction." This methodology first triggers 3D autoregressive modeling by remoulding video diffusion model with unidirectional next-view prediction, where the next-view images are iteratively generated conditioned on the past views and input image. Furthermore, a unique autoregressive scheme via next symmetric view prediction is introduced to simultaneously predict each pair of views under the view-order symmetry at each autoregressive step. Such design captures the view-order symmetry of 360$^{\circ }$ orbital multi-view images and thus encourages strong geometric and appearance consistency among adjacent views in a near-to-far order. Extensive experiments demonstrate that our proposed AR3D achieves strong performance on both novel view synthesis and single view reconstruction tasks for image-to-3D generation.