Chen Yang, Chunyi Chen
To lower technical barriers to content creation in Collaborative Virtual Environments (CVEs), this paper presents “One Word, One World,” a natural-language-driven VR co-creation system. Integrating speech recognition with generative AI, the system allows users to initiate 3D asset generation through spoken commands, receive immediate shared placeholder feedback, and manipulate materialized assets within a multi-user environment supported by soft locking, synchronized state updates, and authorship visualization. We validated the system through a within-subjects study ( N = 48) comparing this voice-driven paradigm against a traditional asset-library baseline. Results indicate that the AI condition significantly increased creative throughput ( p < .001) and expert-rated design quality ( p = .007). Furthermore, the system enhanced social presence and psychological safety, particularly for participants with limited technical backgrounds. These findings suggest that language-driven generative co-creation can improve creative efficiency while lowering technical barriers to visible participation, particularly for users with limited 3D modeling experience.