Green-VLA: Staged Vision-Language-Action Model for Generalist Robots

📅 2026-01-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of generalizing control and ensuring safe execution for general-purpose robots across heterogeneous morphologies—such as humanoids, mobile manipulators, and fixed-base arms—in real-world environments. The authors propose a staged vision-language-action framework that integrates a five-phase curriculum learning strategy, a unified embodied perception-action interface, pretraining with multimodal foundation models, and a safety-enhancement mechanism during inference. Innovatively combining reinforcement learning policy alignment, temporally aligned data processing, and out-of-distribution detection, the approach significantly improves task success rates, robustness, and long-horizon execution efficiency. Extensive experiments in both simulation and real-world robotic platforms demonstrate the method’s effectiveness and performance gains in cross-platform deployment.

Technology Category

Application Category

📝 Abstract
We introduce Green-VLA, a staged Vision-Language-Action (VLA) framework for real-world deployment on the Green humanoid robot while maintaining generalization across diverse embodiments. Green-VLA follows a five stage curriculum: (L0) foundational VLMs, (L1) multimodal grounding, (R0) multi-embodiment pretraining, (R1) embodiment-specific adaptation, and (R2) reinforcement-learning (RL) policy alignment. We couple a scalable data-processing pipeline (3,000 hours of demonstrations) with temporal alignment and quality filtering, and use a unified, embodiment-aware action interface enabling a single policy to control humanoids, mobile manipulators, and fixed-base arms. At inference, the VLA controller is enhanced with episode-progress prediction, out-of-distribution detection, and joint-prediction-based guidance to improve safety and precise target selection. Experiments on Simpler BRIDGE WidowX and CALVIN ABC-D, as well as real-robot evaluations, show strong generalization and performance gains from RL alignment in success rate, robustness, and long-horizon efficiency.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action
generalist robots
embodiment generalization
real-world deployment
robotic control
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language-Action
embodiment-aware policy
staged curriculum learning
reinforcement learning alignment
multimodal grounding
💼 Related Jobs
No related jobs found.
I
I. Apanasevich
Sber Robotics Center
M
M. Artemyev
Sber Robotics Center
R
R. Babakyan
Sber Robotics Center
P
P. Fedotova
Sber Robotics Center
D
D. Grankin
Sber Robotics Center
E
E. Kupryashin
Sber Robotics Center
A
A. Misailidi
Sber Robotics Center
D
D. Nerus
Sber Robotics Center
A
A. Nutalapati
Sber Robotics Center
G
G. Sidorov
Sber Robotics Center
I
I. Efremov
Sber Robotics Center
M
M. Gerasyov
Sber Robotics Center
D
D. Pikurov
Sber Robotics Center
Y
Y. Senchenko
Sber Robotics Center
S
S. Davidenko
Sber Robotics Center
D
D. Kulikov
Sber Robotics Center
M
M. Sultankin
Sber Robotics Center
K
K. Askarbek
Sber Robotics Center
O
O. Shamanin
Sber Robotics Center
D
D. Statovoy
Sber Robotics Center
E
E. Zalyaev
Sber Robotics Center
I
I. Zorin
Sber Robotics Center
A
A. Letkin
Sber Robotics Center
E
E. Rusakov
Sber Robotics Center
A
A. Silchenko
Sber Robotics Center