VLA-ULAP: Interleaving Cloud VLA Calls with Ultra-Lightweight Local Action Prediction at the Edge

📅 2026-09-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决大规模视觉-语言-动作模型在边缘设备上的高能耗和延迟问题,提出VLA-ULAP方法,通过超轻量级本地预测减少远程调用,提高响应速度并降低能耗。
📝 Abstract
Billion-parameter vision--language--action (VLA) policies demand substantial onboard power, while communication delays in remote inference hinder timely responses. We propose VLA-ULAP, which interleaves remote VLA calls with an Ultra-Lightweight Local Action Predictor (ULAP). With approximately 7.4M parameters including the frozen vision encoder, ULAP combines current views, proprioception, and executed action history to predict chunks in one pass. Trained independently, it requires no VLA hidden states, online verification, or server round trips. On Jetson Orin Nano, ULAP takes 19.9 ms and 0.183 J per inference, compared with 284.3 ms and 50.55 J for GR00T on RTX A6000. Across three simulated base-policy/benchmark pairs, selected operating points remove 48.8--76.7\% of VLA calls while retaining 95.0--97.5\% of the baseline success rate. Against local VLA-acceleration alternatives on VLA-JEPA, ULAP uses an estimated 49.2\% less inference time and 51.0\% less GPU energy per successful episode than ACT at comparable success rates, and 77.1\% less time and 79.9\% less energy than SP-VLA at equal success rates. Physical SO-101 experiments retain 95.2--100\% of the baseline success rate across seen and held-out placements while reducing inference time by an estimated 47.9--58.0\% and inference-device energy by 52.1--62.5\%, based on successful-episode call counts and measured device costs. Faster responses also improve dynamic-task success rates: in latency-aware LIBERO-Safety simulation, VLA-ULAP exceeds $π_{0.5}$ by 11.0 and 15.5 percentage points on two tasks while approximately halving VLA calls.
Problem

Research questions and friction points this paper is trying to address.

VLA
communication delay
onboard power
local action prediction
energy efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

VLA-ULAP
Ultra-Lightweight Local Action Predictor (ULAP)
edge computing
energy efficiency
remote inference
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Deyu Cao
Deyu Cao
the University of Tokyo, University of Toronto
R
Ryuji Oi
Institute of Science Tokyo
K
Kosuke Matsushima
Institute of Science Tokyo
Y
Yuxuan Pan
The University of Tokyo
Z
Ziheng Wang
The University of Tokyo
Daichi Fujiki
Daichi Fujiki
Institute of Science Tokyo
Computer architecture
A
Atsutake Kosuge
The University of Tokyo