Knowing When to Stop: Adaptive Action Chunking via Internal Cross-Attention Dynamics in VLAs

๐Ÿ“… 2026-09-01
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
ๆœฌๆ–‡ๆๅ‡บไธ€็งๅŸบไบŽๅ†…้ƒจไบคๅ‰ๆณจๆ„ๅŠ›ๅŠจๆ€็š„่‡ช้€‚ๅบ”ๅŠจไฝœๅˆ†ๅ—ๆ–นๆณ•๏ผŒไปฅ่งฃๅ†ณๅ›บๅฎšๆ‰ง่กŒ่Œƒๅ›ดๅœจๆ•ˆ็އๅ’Œๅ‡†็กฎๆ€งไน‹้—ด็š„ๆƒ่กก้—ฎ้ข˜ใ€‚
๐Ÿ“ Abstract
Action chunking is a standard execution strategy in modern Vision-Language-Action (VLA) frameworks, but fixed execution horizons impose a trade-off between efficiency and accuracy. Short chunks require frequent inference and may cause oscillatory behavior, whereas long chunks can become misaligned with newly observed states. We address this limitation with an adaptive action chunking approach based on internal cross-attention dynamics in the action expert. We observe that, as the prediction horizon extends, action-to-observation cross-attention becomes increasingly dispersed and its entropy rises toward a plateau. This pattern is associated with higher action prediction error and provides an online signal that the current observation offers limited grounding for further open-loop execution. Based on this observation, we introduce a training-free truncation mechanism that detects sustained high-entropy plateaus and dynamically selects the execution horizon during inference. The method uses attention weights already computed by the policy and introduces negligible additional overhead. Evaluations on $ฯ€_{0.5}$ and X-VLA across RoboTwin 2.0, LIBERO, and three real-world manipulation tasks show improved average task success over fixed-horizon and adaptive chunking baselines, while preserving efficient closed-loop control. These results show that cross-attention dynamics can provide a practical internal signal for adaptive action execution in VLAs.
Problem

Research questions and friction points this paper is trying to address.

action chunking
execution horizon
cross-attention dynamics
adaptive action execution
Innovation

Methods, ideas, or system contributions that make the work stand out.

adaptive action chunking
cross-attention dynamics
entropy plateau
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
R
Runze Xu
Tsinghua University
X
Xiaolong Shan
Tsinghua University
S
Shuang Dai
Tsinghua University
Y
Yu Wang
Tsinghua University
Jincheng Yu
Jincheng Yu
Tsinghua University
FPGARobotics