A Survey on Vision-Language-Action Models for Embodied AI
This paper addresses the core challenge of how Vision-Language-Action (VLA) models support language-conditioned robotic tasks in embodied intelligence. Methodologically, it introduces the first systematic, panoramic survey framework, proposing a three-dimensional taxonomy—“Component Design–Low-level Action Policies–High-level Task Planning”—that unifies VLA modeling, embodied control, task decomposition, simulation integration, and cross-benchmark evaluation. Key contributions include: (1) the first explicit characterization of three principal VLA technical paradigms; (2) a comprehensive survey of multimodal datasets, embodied simulation platforms, and standardized evaluation benchmarks; and (3) a structured knowledge graph that identifies critical open challenges—including scalable architecture design, world model integration, and real-world deployment—and outlines promising future research directions.