ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models

📅 2026-09-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出ActionPiece,通过联合监督表示学习和量化来保持物理动作关系,以解决自回归视觉-语言-动作模型中动作标记化保真度问题。
📝 Abstract
Action tokenizers play a central role in autoregressive vision-language-action (VLA) models, determining both the targets for policy training and the executable commands recovered from predicted tokens. Their fidelity is commonly evaluated using pointwise reconstruction metrics such as mean squared error (MSE), yet small individual errors do not fully characterize how faithfully action adjustments across demonstrations are preserved. After compression, similar actions may still cluster around a representative motion, while the adjustments needed for different contexts are diminished, distorted, or even reversed. We introduce physical rank consistency (PRC) to measure how well tokenization preserves local physical distance rankings after reconstruction. Evaluating decoded actions provides a common reference across token vocabularies and decoder architectures, complementing pointwise accuracy with a measure of relational fidelity. We further present ActionPiece, which preserves physical action relationships through joint supervision of representation learning and quantization. Physical rank preservation supervises near-far ordering in encoder and quantized feature distances, while quantization regularization applies the same ordering to codeword assignment distributions. Both objectives augment reconstruction, producing discrete action tokens for standard autoregressive policy learning and execution through a frozen decoder. Under the same Qwen3-VL-4B policy training setup, ActionPiece achieves 94.8% on LIBERO and 68.8% on unseen LIBERO-Plus, with additional evaluations reaching 71.9% on SimplerEnv and 51.5% across VLA-Arena L0-L2. Component ablations show that the two objectives jointly improve PRC and policy success, demonstrating the value of physical relationship supervision for action tokenization.
Problem

Research questions and friction points this paper is trying to address.

Action Tokenization
Reconstruction Metrics
Physical Rank Consistency
Autoregressive Vision-Language-Action Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

physical rank consistency
action tokenization
representation learning
quantization
autoregressive vision-language-action models
🔎 Similar Papers
No similar papers found.
S
Shijie Lian
Huazhong University of Science and Technology
B
Bin Yu
Harbin Institute of Technology
Z
Zhaolong Shen
Beihang University
X
Xiaopeng Lin
The Hong Kong University of Science and Technology (Guangzhou)
Y
Yichao Du
DeepCybo
Z
Zhirui Zhang
DeepCybo
L
Laurence T. Yang
Huazhong University of Science and Technology
K
Kai Chen
Zhongguancun Institute of Artificial Intelligence