Sensorimotor features of self-awareness in multimodal large language models
This study investigates whether multimodal large language models (MLLMs) can spontaneously develop bodily self-awareness solely through embodied sensorimotor interaction. We embed an MLLM in an autonomous mobile robot that explores its environment and learns closed-loop behaviors exclusively from real-time multimodal sensory inputs—vision, touch, proprioception, and vestibular signals—without any explicit supervision or pre-defined self-models. We systematically evaluate the model’s capabilities in environmental recognition, self-discrimination, and motor prediction. Our key contributions are threefold: (1) First empirical evidence that MLLMs hierarchically emergent bodily self-awareness in a fully unsupervised, embodied setting; (2) Causal insights—derived via structural equation modeling and sensory ablation experiments—into how multisensory integration, temporal memory, and hierarchical internal representations jointly enable self-awareness; and (3) Demonstration that structured and episodic memory are essential for coherent self-referential reasoning, along with identification of critical sensory modalities and their functional redundancy relationships.