Green-VLA: Staged Vision-Language-Action Model for Generalist Robots
This work addresses the challenge of generalizing control and ensuring safe execution for general-purpose robots across heterogeneous morphologies—such as humanoids, mobile manipulators, and fixed-base arms—in real-world environments. The authors propose a staged vision-language-action framework that integrates a five-phase curriculum learning strategy, a unified embodied perception-action interface, pretraining with multimodal foundation models, and a safety-enhancement mechanism during inference. Innovatively combining reinforcement learning policy alignment, temporally aligned data processing, and out-of-distribution detection, the approach significantly improves task success rates, robustness, and long-horizon execution efficiency. Extensive experiments in both simulation and real-world robotic platforms demonstrate the method’s effectiveness and performance gains in cross-platform deployment.