🤖 AI Summary
This study addresses critical bottlenecks in Vision-Language-Action (VLA) model deployment, including the efficiency-performance trade-off, limited cross-embodiment generalization, and jerky motion generation. To overcome these challenges, we propose an asynchronous dual-frequency architecture that decouples semantic reasoning from action control. Furthermore, we introduce GESTURE-7, a unified action representation, alongside a mask-guided smoothing constraint algorithm. Experimental evaluations on the LIBERO-Plus benchmark demonstrate that our method achieves an average success rate of 85.5% and accelerates action generation by approximately 2.7× compared to synchronous baselines. These results confirm that the proposed framework effectively enables efficient, smooth, and generalizable robotic manipulation, significantly outperforming existing approaches in both computational efficiency and operational robustness.
📝 Abstract
Real-world deployment of Vision-Language-Action (VLA) models is often bottlenecked by efficiency-performance trade-offs, cross-embodiment generalization, and execution smoothness. We present NebulaVLA, an asynchronous dual-frequency architecture that decouples high-level semantic reasoning from low-level action control, optimizing computational resources and modularity. To bridge semantic gaps across heterogeneous robots, we introduce GESTURE-7, a unified language-grounded action representation. Furthermore, our Guide Action algorithm enforces kinematic continuity via mask-based smoothness constraints. Comprehensive evaluations demonstrate that NebulaVLA significantly outperforms synchronous baselines, achieving an 85.5\% average success rate on LIBERO-Plus and accelerating action generation by \textasciitilde 2.7$\times$. This asynchronous design enables highly efficient and responsive control for practical robotics.