MobileVLA-R1 2.0: RL-Enhanced Reasoning for Mobile Robot Control

📅 2026-09-05
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决移动机器人中自然语言指令转化为可靠执行动作的问题,提出MobileVLA-R1 2.0框架,结合强化学习增强的结构化推理与控制,提升长时决策一致性。
📝 Abstract
Grounding natural-language instructions into reliable and executable actions remains a fundamental challenge for vision-language-action (VLA) systems on mobile robots, due to the persistent gap between high-level semantic reasoning and low-level locomotion and manipulation control. Existing approaches often rely on implicit reasoning or monolithic action prediction, making it difficult to maintain coherent long-horizon decision making while producing precise and adaptable robot actions. To address this challenge, we propose MobileVLA-R1 2.0, an RL-enhanced VLA framework that explicitly couples structured embodied reasoning with executable mobile robot control. The framework learns multi-granularity reasoning over embodied trajectories through supervised Chain-of-Thought (CoT) alignment and reinforcement learning, improving reasoning-to-action consistency beyond purely behavioral supervision. To support both locomotion and manipulation, we further introduce a reasoning-conditioned action decoder that maps multimodal reasoning representations to task-level action targets, which are subsequently translated into embodiment-specific commands by robot controllers. This design provides a unified perception-reasoning-action interface while decoupling high-level action generation from robot-specific actuation. We conduct extensive evaluations on language-guided navigation, quadruped control, and humanoid mobile manipulation, covering VLN-CE, QUARD, and real-world deployments on Unitree Go2 and G1 robots. MobileVLA-R1 2.0 consistently outperforms strong VLA baselines, achieving an average 1.6 point improvement in SR on VLN-CE and a 10.0 point improvement in full-task success on real-world G1 mobile manipulation tasks over MobileVLA-R1, while demonstrating robust long-horizon instruction following and closed-loop execution across different robotic platforms.
Problem

Research questions and friction points this paper is trying to address.

natural-language instructions
visual-language-action systems
mobile robots
semantic reasoning
locomotion and manipulation control
Innovation

Methods, ideas, or system contributions that make the work stand out.

RL-Enhanced VLA Framework
Chain-of-Thought (CoT) Alignment
Reasoning-Conditioned Action Decoder
Multi-Granularity Reasoning
Unified Perception-Reasoning-Action Interface
🔎 Similar Papers
No similar papers found.
T
Ting Huang
School of Computer Science, Peking University, Beijing 100871, China
Y
Yue Huang
South China University of Technology, Guangzhou 510006, China
Zeyu Zhang
Zeyu Zhang
Gaoling School of Artificial Intelligence, Renmin University of China
LLM-based AgentResponsible RecSysCausal Learning
S
Shuicheng Yan
School of Computing, National University of Singapore, Singapore 117417
Hao Tang
Hao Tang
Peking University
computer vision