🤖 AI Summary
This work addresses the computational inefficiency in Vision-Language-Action (VLA) models during real-time control, where redundant recomputation of nearly identical visual tokens across consecutive frames leads to significant waste in KV cache usage. Existing caching strategies rely on heuristic visual similarity metrics and overlook the model’s internal uncertainty, compromising reliability. To overcome this, the authors propose Gated VLA-Cache, which introduces, for the first time, the logit margin from action token predictions as a neural introspection signal to dynamically determine KV cache reuse—without requiring any additional training. This approach enables adaptive, zero-overhead cache gating that fully recovers or even surpasses the original accuracy on LIBERO-Goal and LIBERO-Long benchmarks while preserving 80% computational savings, substantially enhancing both efficiency and reliability across all four LIBERO suites.
📝 Abstract
Vision-Language-Action(VLA) models map camera images and language instructions directly to motor commands through a single autoregressive transformer. In real-time control, they still spend substantial compute recomputing key-value(KV) representations for visual tokens that barely change across neighboring frames. Recent work such as VLA-Cache reduces that cost by reusing KV states for visually static patches, but its policy relies only on observation-space heuristics and does not account for the model's own uncertainty. We propose Gated VLA-Cache, a lightweight, training-free extension that augments visual-similarity caching with neural introspection. The method monitors the logit margin between the top two predicted action tokens, a zero-cost confidence signal available during decoding. When the margin drops below a threshold, the cache is invalidated and a full recompute is triggered. Evaluated on four LIBERO benchmark suites with both OpenVLA and OpenVLA-OFT, Gated VLA-Cache improves reliability when blind caching hurts. On LIBERO-Goal and LIBERO-Long, it recovers over 100% of the lost accuracy while retaining 80% of the compute savings.