World-to-Wrist: Task-Conditioned Future Wrist Modeling for Fine-Grained Robot Manipulation

📅 2026-08-05
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of existing vision-language-action (VLA) models, which fail to differentiate the roles of egocentric and wrist-mounted views in fine manipulation and lack explicit modeling of how wrist interactions evolve within task contexts. To bridge this gap, we propose W2-VLA, a novel framework that introduces, for the first time, a task-conditioned future wrist latent prediction mechanism to establish a compact interface between global task context and local wrist interaction. We further design a W2-CoT synthesis pipeline that generates structured supervision signals incorporating manipulation progress, physical transition cues, and wrist-based evidence, thereby enhancing semantic alignment of the latent variables. Experiments demonstrate that our approach significantly improves fine-grained and contact-sensitive manipulation performance for both single- and dual-arm systems on LIBERO, RoboTwin 2.0, and real-world tasks, achieving action generation rates exceeding 80 Hz.
📝 Abstract
Vision-language-action (VLA) models often treat main-view and wrist-view observations as parallel visual inputs, overlooking their distinct roles in robot manipulation. Fine-grained manipulation, however, benefits from anticipating how wrist-local interactions may evolve under the global task context. To address this limitation, we present World-to-Wrist VLA (W2-VLA), a VLA model for fine-grained robot manipulation with task-conditioned future wrist modeling. Given current multi-view observations and a task instruction, W2-VLA contextualizes a set of latent modeling tokens as a compact interface between the vision-language model and the wrist predictor. Conditioned on this interface and the observed wrist history, the predictor forecasts future wrist latents, which are transformed into future-aware context for action prediction. In addition, we introduce W2-CoT, a synthesis pipeline that produces structured annotations describing manipulation progress, physical transition cues, and wrist-local evidence. These annotations provide auxiliary supervision that shapes the task-conditioned latent interface. Experiments on LIBERO, RoboTwin 2.0, and real-world manipulation tasks demonstrate improved fine-grained and contact-sensitive manipulation across both single-arm and bimanual settings, while maintaining action-generation rates above 80 Hz.
Problem

Research questions and friction points this paper is trying to address.

fine-grained manipulation
vision-language-action models
wrist-view prediction
task-conditioned modeling
robot manipulation
Innovation

Methods, ideas, or system contributions that make the work stand out.

task-conditioned wrist modeling
vision-language-action (VLA)
future-aware manipulation
latent interface
fine-grained robot manipulation
💼 Related Jobs
No related jobs found.
Y
Yuhao Pan
The Hong Kong University of Science and Technology
H
Haosong Peng
The Hong Kong University of Science and Technology
Z
Zhengshen Zhang
National University of Singapore
Z
Zhengyang Yan
The Hong Kong University of Science and Technology
Yalun Dai
Yalun Dai
Nanyang Technological University
deep learning
Fushuo Huo
Fushuo Huo
The Hong Kong Polytechnic University
Large Vision Language ModelMultimodal LearningTrustworthy AI
C
Chujie Wang
Wuhan University
T
Tianyu Qi
Sun Yat-sen University
Xiucheng Wang
Xiucheng Wang
Xidian University
wireless communicationgraph neural networkreinforcement learningdigital twin
Nan Cheng
Nan Cheng
University of Michigan
condensed matter physics
Wenchao Xu
Wenchao Xu
Hong Kong University of Science and Technology
Multimodal learningdistributed computingInternet of things