Takazumi Matsumoto, Prasanna Vijayaraghavan, Jun Tani
This review presents goal-directed latent variable inference (GLean) as a unified active inference framework for embodied intelligence operating in continuous spatiotemporal domains. We address a central challenge in cognitive science: how agents generate, adapt, and understand goal-directed actions in noisy, high-dimensional sensorimotor environments while flexibly representing goals as static states, dynamic trajectories, or abstract linguistic constraints. GLean formulates planning as inference in a hierarchical latent space, enabling efficient real-time action generation without explicit policy sampling. We first introduce T-GLean, which integrates teleological goal representation with continuous latent-variable planning, demonstrating robust online plan adaptation to dynamically changing environments in humanoid robot experiments. We then review extensions incorporating compositional language grounding through multimodal generative modeling, enabling generalization to novel verb-noun combinations in goal representation. Together, this body of work bridges active inference, cognitive neuroscience, and embodied robotics, offering a scalable computational account of flexible, teleological action in continuous domains.