GWM-VLA: Geometry-Aware Latent World Modeling for Vision-Language-Action Learning

📅 2026-08-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limited robustness of existing vision-language-action (VLA) models under environmental and viewpoint variations, as well as their neglect of multi-view geometric relationships. To overcome these limitations, the authors propose a geometry-aware implicit world model that leverages a VGGT-Ω module to encode multi-view states and predict local tokens for a target viewpoint, thereby preserving critical geometric information. The approach further introduces a shared implicit action representation to unify learning signals between the world model and the action head, integrating a flow-matching action head with supervision from real robot actions. Evaluated in both simulation and real-world settings, the method significantly enhances manipulation performance, particularly excelling in modeling end-effector–object interactions from wrist-mounted camera views.
📝 Abstract
Vision-Language-Action (VLA) models achieve strong robotic manipulation performance but often degrade under visual and environmental shifts. Latent world modeling offers a promising approach to improving robustness, yet existing methods commonly encode camera views independently and predict holistic scene dynamics without explicitly modeling their geometric relationships. We propose GWM-VLA, a geometry-aware latent world modeling framework for VLA learning. GWM-VLA combines geometry-aware multi-view state encoding, global context-conditioned target-view prediction, and shared latent-action representations grounded by robot-action supervision. Specifically, VGGT-$Ω$ jointly aggregates multi-view observations at each timestep to construct geometry-aware multi-view states. The latent world model predicts the next-step patch tokens of a selected target view using patch and register tokens obtained after multi-view aggregation, thereby retaining multi-view geometric information without predicting the complete multi-view state. We use the wrist view as the target in our experiments, placing greater emphasis on end-effector motion and local gripper-object interactions. Finally, the shared latent-action representations condition both the latent world model and the flow-matching action head, allowing latent-prediction supervision and ground-truth robot-action supervision to jointly shape the same latent-action representations. Experiments across both simulation and real-world environments demonstrate the effectiveness and robustness of GWM-VLA.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action
Latent World Modeling
Geometric Relationships
Robustness
Visual Shifts
Innovation

Methods, ideas, or system contributions that make the work stand out.

geometry-aware modeling
latent world model
multi-view state encoding
shared latent-action representation
vision-language-action learning
🔎 Similar Papers
No similar papers found.