科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Neuron2026-04-01· Situated

Embodiment in multimodal large language models

Akila Kadambi, Lisa Aziz-Zadeh, António R. Damásio, Marco Iacoboni, Srini Narayanan

原始摘要(英文原文)· Original abstract
Multimodal large language models (MLLMs) have demonstrated an extraordinary capacity to bridge textual and visual inputs. Nonetheless, MLLMs still face limitations in situated physical and social interactions in sensorially rich and multimodal real-world settings, where the embodied experience of a living organism appears fundamental. We suggest that the next frontiers for MLLM development require the incorporation of both internal and external embodiment-modeling not only external interactions with the world but also internal states and drives. Here, we describe mechanisms of internal and external embodiment in humans and relate these to current advances in MLLMs in the early stages of aligning to human representations. Our dual-embodied framework proposes to model interactions between these forms of embodiment in MLLMs so as to bridge the gap between multimodal data and world experience.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Embodiment in multimodal large language models — 科研速览 Science Skim