Tadahiro Taniguchi, Ryo Ueda, Tomoaki Nakamura, Masahiro Suzuki, Akira Taniguchi
This survey and perspective paper addresses a fundamental puzzle: how Large Language Models (LLMs) acquire extensive world knowledge without direct sensorimotor experience. We introduce the Collective World Model hypothesis, arguing that an LLM does not learn a world model from scratch. Instead, it learns a statistical approximation of a collective world model already implicitly encoded in human language through a society-wide process of embodied, interactive sense-making. To formalize this process, we introduce the generative emergent communication (Generative EmCom) framework, built upon Collective Predictive Coding (CPC). This framework models language emergence as decentralized Bayesian inference, effectively creating an encoder-decoder structure at a societal scale: human society collectively encodes its grounded representations into language, and an LLM decodes these symbols to reconstruct a latent space mirroring the original structure. This perspective offers a principled explanation for how LLMs acquire their capabilities. We formalize the Generative EmCom framework, connecting it to world models and multi-agent reinforcement learning, and apply it to interpret LLMs, explaining phenomena like distributional semantics as a natural consequence of representation reconstruction. This work provides a unified theory that bridges individual cognitive development, collective language evolution, and the foundations of large-scale AI.