Algorithm-Architecture Co-Design for Efficient VLA Inference via Speculative Inference and Verification

📅 2026-08-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the real-time deployment challenges of Vision-Language-Action (VLA) models caused by high computational costs and constrained action prediction. We propose SpecVLA, an algorithm-system co-design framework that introduces an environment-aware speculative inference paradigm. By constructing a hardware-friendly verification model and integrating differential residuals, mixed-precision quantization, and GPU-specific heterogeneous parallel dataflows, SpecVLA enables efficient speculation and verification. Evaluations on the LIBERO and ManiSkill benchmarks demonstrate that SpecVLA significantly reduces end-to-end latency while maintaining task success rates. Consequently, this framework effectively facilitates efficient and reliable real-time robotic manipulation, overcoming critical bottlenecks in deploying large-scale VLA models for practical applications.
📝 Abstract
Vision-Language-Action (VLA) models have demonstrated remarkable capabilities in the field of embodied AI, but their high computational cost and limited predicted action length hinder real-time deployment. Although Dadu-Corki, a dedicated accelerator for efficient embodied AI, has been introduced, it does not exploit the inherent interaction patterns between the robot and its environment, which results in a relatively short predicted action length. We observe that robotic environments naturally alternate between active states-where precise actions are crucial-and inactive states-where actions have limited impact on task success. This insight enables a new scheduling opportunity: long-action-length speculative prediction in inactive states, paired with selective verification in active states. We propose SpecVLA, an algorithm-system co-design framework that adaptively balances action length, inference latency, and task reliability. On the algorithm side, SpecVLA introduces a state-aware VLA inference execution paradigm and a hardware-friendly construction of a smaller verification model (sVLA) using differential residuals and block-wise mixed-precision quantization. On the system side, we develop a heterogeneous architecture consisting of a GPU and a robotic-specific hardware module, along with a speculative dataflow that decouples VLA and sVLA through parallel execution. Comprehensive evaluations on OpenVLA and RDT across LIBERO and ManiSkill benchmarks show that SpecVLA reduces end-to-end latency significantly while preserving task success rate. By enabling long-action-length speculative prediction with timely verification, SpecVLA achieves real-time robotic manipulation with both high efficiency and reliability.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action
Embodied AI
Real-time Inference
Action Prediction Length
Computational Efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Algorithm-Architecture Co-Design
Speculative Inference
Vision-Language-Action Models
State-Aware Scheduling
Heterogeneous Architecture
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
C
Chunyu Qi
School of Computer Science, Shanghai Jiao Tong University
Z
Zhuoran Song
School of Computer Science, Shanghai Jiao Tong University
Jian Weng
Jian Weng
King Abdullah University of Science and Technology
Computer ArchitectureCompiler Optimization
Haozhe Jiang
Haozhe Jiang
PhD Student, EECS, UCBerkeley
Artificial Intelligence
X
Xueyuan Liu
School of Computer Science, Shanghai Jiao Tong University
N
Naifeng Jing
School of Computer Science, Shanghai Jiao Tong University
G
Guanghui He
School of Computer Science, Shanghai Jiao Tong University
Xiaoyao Liang
Xiaoyao Liang
Shanghai Jiao Tong University
Computer Architecture
H
Haibing Guan
School of Computer Science, Shanghai Jiao Tong University