Yan Zhou, Jinping Zhang, Bangzhen Liu, Dewang Ye, Qinghua Lu, Xuemiao Xu, Cheng Xu, Yuexia Zhou
Zero-shot text-to-3D generation has garnered growing interest through the integration of vision-language foundation models and 3D generative frameworks. However, existing methods often overlook the preservation of local geometric semantics critical for producing semantically aligned and structurally coherent 3D point clouds. Here, we introduce RecPoint, a structure-aware framework for text-guided 3D shape generation that rectifies semantic drift by explicitly modeling local geometric relationships. Central to our method is a geometry-aware graph transformer that encodes point-wise adjacency into graph signals, enabling fine-grained feature learning and robust shape reconstruction. Complementing this, we propose a test-time locality-aware semantic rectification mechanism that bridges modality gaps by grounding textual features in visually similar, structurally aligned image embeddings. Without relying on paired text-shape data, RecPoint achieves superior fidelity and semantic alignment across standard benchmarks, outperforming previous state-of-the-art approaches.