$τ_0$-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation

📅 2026-08-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of critical decision-making compute allocation mechanisms in hierarchical models for long-horizon manipulation by proposing τ0-VLA. This method introduces test-time compute into hierarchical Vision-Language-Action architectures for the first time, leveraging world models to guide reasoning and execution memory. By reformulating high-level subtask generation as a scalable dynamic search problem, it achieves decoupling between high- and low-level policies while enabling multimodal collaborative training. Experimental results demonstrate that this mechanism significantly improves both subtask prediction accuracy and closed-loop success rates in long-horizon tasks. These findings validate the effectiveness of on-demand compute allocation in enhancing robotic decision-making capabilities for complex, extended operations.
📝 Abstract
Long-horizon robot manipulation requires a robot to both execute individual skills reliably and sequence them coherently over extended tasks. Most hierarchical vision-language-action (VLA) models make each such decision with a single forward pass, leaving no mechanism to allocate additional computation to difficult or consequential choices. We introduce $τ_0$-VLA, a hierarchical robot foundation model that formulates high-level subtask generation as a compute-scalable inference problem through world-model-guided test-time computation. At each inference step, the high-level policy uses execution memory to generate a subtask and, when needed, searches over alternatives before committing to its output. A low-level policy then executes the generated subtask across multiple robot embodiments. The policy is trained on 40,115 hours of heterogeneous real-world data with multimodal co-training. Across in-domain and distribution-shifted settings, allocating additional test-time computation substantially improves next-subtask prediction accuracy, and these gains translate into higher closed-loop success on long-horizon robot manipulation tasks.
Problem

Research questions and friction points this paper is trying to address.

Long-horizon robot manipulation
Hierarchical VLA models
Test-time computation
Subtask generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Test-Time Computation
Hierarchical VLA
World Model
Compute-Scalable Inference
Long-Horizon Manipulation
💼 Related Jobs
No related jobs found.
X
Xiaowei Cai
Agibot Finch
Y
Yunuo Cai
Shanghai Innovation Institute, Agibot Finch
B
Bingao Chen
Agibot Finch, The Chinese University of Hong Kong
J
Jingxiao Chen
Agibot Finch
Z
Zhi Chen
Agibot Finch
S
Siyuan Feng
Agibot Finch
Tengyu Hou
Tengyu Hou
Agibot Finch
J
Jingshun Huang
Shanghai Innovation Institute, Agibot Finch
Han Jiang
Han Jiang
Johns Hopkins University
Natural Language GenerationSocietal AIModel Evaluation
R
Runkun Ju
Agibot Finch
D
Dong Li
Agibot Finch
M
Mingxiang Li
Agibot Finch
Shaowei Li
Shaowei Li
University of California, San Diego
Chemical Physics and Physical Chemistry
X
Xinchen Li
Agibot Finch
Y
Yifan Li
Shanghai Innovation Institute, Agibot Finch
Y
Yi Liu
Shanghai Innovation Institute, Agibot Finch
Zhongyuan Liu
Zhongyuan Liu
Tencent
AIGC Games
Jianlan Luo
Jianlan Luo
UC Berkeley, Google X
RoboticsMachine LearningArtificial Intelligence
J
Junwen Miao
Agibot Finch
Ruiqi Ni
Ruiqi Ni
Purdue University
RoboticsComputer Graphics
Buqing Nie
Buqing Nie
Shanghai Jiao Tong University
Reinforcement LearningRobot Learning
Mingjie Pan
Mingjie Pan
Peking University
X
Xinlin Ren
Agibot Finch
J
Jianheng Song
Agibot Finch
J
Jiaxu Wang
Agibot Finch, The Chinese University of Hong Kong