TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM

šŸ“… 2026-07-29
šŸ“ˆ Citations: 0
✨ Influential: 0
šŸ“„ PDF
šŸ¤– AI Summary
This work proposes TurboVLA, an efficient vision-language-action (VLA) framework that eliminates the reliance on large language models (LLMs) as a central hub between perception and action, thereby significantly reducing computational overhead and GPU memory consumption. Instead of LLM-centric architectures, TurboVLA establishes an end-to-end V+L→A mapping by independently encoding visual and linguistic inputs, fusing them through a lightweight bidirectional interaction module, and predicting continuous actions via a compact decoder. With only 0.2 billion parameters, TurboVLA achieves an average success rate of 97.7% on the LIBERO benchmark, while maintaining a low inference latency of 31.2 ms and consuming merely 0.9 GB of GPU memory. On an RTX 4090, it delivers real-time performance at 32 Hz, matching or surpassing substantially larger models in both efficiency and effectiveness.
šŸ“ Abstract
Vision-language-action (VLA) models commonly adopt an LLM-centric $V \to L \to A$ pathway, where visual observations are projected into the representation space of a large language model before being decoded into robot actions. Although effective, this design incurs substantial computation and memory overhead at every policy invocation. In this work, we introduce TurboVLA, a new VLA paradigm that reformulates the conventional $V \to L \to A$ pathway as a direct $V + L \to A$ mapping. Instead of using a large language model as the central interface between perception and action, TurboVLA independently encodes visual observations and language instructions, directly exchanges information between them through lightweight bidirectional vision-language interaction, and predicts continuous action chunks with a compact decoder. This simple design constructs task-conditioned representations directly from visual and linguistic features, significantly reducing the computational and memory costs of VLA inference. On LIBERO, TurboVLA achieves 97.7% average success with only 0.2B parameters, 31.2 ms inference latency, and 0.9 GB inference VRAM on a consumer-grade RTX 4090, matching or outperforming substantially larger VLA policies. These results establish TurboVLA as a simple and effective alternative to the prevailing LLM-centric VLA paradigm, offering a new perspective on how vision, language, and action can be connected for efficient robotic manipulation. Code is available at https://github.com/H-EmbodVis/TurboVLA.
Problem

Research questions and friction points this paper is trying to address.

vision-language-action
computational overhead
memory efficiency
real-time inference
robotic manipulation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language-Action
Efficient Robotics
Direct V+L to A Mapping
Lightweight Interaction
Real-Time Inference
šŸ’¼ Related Jobs
No related jobs found.
H
Hengyi Xie
Huazhong University of Science and Technology
C
Chenfei Yao
Huazhong University of Science and Technology
X
Xianjin Wu
Huazhong University of Science and Technology
X
Xuanyang Xi
Huawei Technologies Co. Ltd, China
Y
Yiping Tang
Huawei Technologies Co. Ltd, China
Di Xu
Di Xu
Professor
Economics of EducationHigher Education PolicyProgram EvaluationCommunity CollegesOnline
Yingying Zhu
Yingying Zhu
PhD student, Dept. of EI, Huazhong University of Science and Technology
computer visionmachine learning
Dingkang Liang
Dingkang Liang
Huazhong University of Science and Technology
Embodied AIWorld ModelAutonomous DrivingCrowd Counting
Xiang Bai
Xiang Bai
Huazhong University of Science and Technology (HUST)
Computer VisionOCR
H
Han Ding
Huazhong University of Science and Technology