NebulaVLA: A Dual-Frequency Vision-Language-Action Model With Guide Action for Robotic Manipulation

📅 2026-08-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses critical bottlenecks in Vision-Language-Action (VLA) model deployment, including the efficiency-performance trade-off, limited cross-embodiment generalization, and jerky motion generation. To overcome these challenges, we propose an asynchronous dual-frequency architecture that decouples semantic reasoning from action control. Furthermore, we introduce GESTURE-7, a unified action representation, alongside a mask-guided smoothing constraint algorithm. Experimental evaluations on the LIBERO-Plus benchmark demonstrate that our method achieves an average success rate of 85.5% and accelerates action generation by approximately 2.7× compared to synchronous baselines. These results confirm that the proposed framework effectively enables efficient, smooth, and generalizable robotic manipulation, significantly outperforming existing approaches in both computational efficiency and operational robustness.
📝 Abstract
Real-world deployment of Vision-Language-Action (VLA) models is often bottlenecked by efficiency-performance trade-offs, cross-embodiment generalization, and execution smoothness. We present NebulaVLA, an asynchronous dual-frequency architecture that decouples high-level semantic reasoning from low-level action control, optimizing computational resources and modularity. To bridge semantic gaps across heterogeneous robots, we introduce GESTURE-7, a unified language-grounded action representation. Furthermore, our Guide Action algorithm enforces kinematic continuity via mask-based smoothness constraints. Comprehensive evaluations demonstrate that NebulaVLA significantly outperforms synchronous baselines, achieving an 85.5\% average success rate on LIBERO-Plus and accelerating action generation by \textasciitilde 2.7$\times$. This asynchronous design enables highly efficient and responsive control for practical robotics.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action
Robotic Manipulation
Cross-embodiment Generalization
Execution Smoothness
Efficiency-performance Trade-offs
Innovation

Methods, ideas, or system contributions that make the work stand out.

Asynchronous Dual-Frequency Architecture
GESTURE-7
Guide Action
Vision-Language-Action Model
Cross-Embodiment Generalization
💼 Related Jobs
No related jobs found.
C
Cong Zhao
ZTE Corporation
S
Shuai Tian
ZTE Corporation
X
Xu Zhang
ZTE Corporation
B
Baocheng Ni
ZTE Corporation
X
Xinguo Song
ZTE Corporation
X
Xueying Sun
ZTE Corporation
S
Shu Jiang
ZTE Corporation
S
Shouchang Yang
ZTE Corporation
B
Bo Tang
ZTE Corporation
J
Jin Deng
ZTE Corporation
Ge Zhu
Ge Zhu
Adobe Research, Music AI
Audio UnderstandingAudio Generative Models
Y
YongCheng Wang
ZTE Corporation
J
Jin Xu
ZTE Corporation
R
Ri Yang
ZTE Corporation