V-Link: Recovering Lost Visual Representations in Action DiT for Vision-Language-Action Models

📅 2026-08-25
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对VLA模型中动作专家难以获取视觉几何和语义信息的问题,提出V-Link方法恢复这些表示,提高机器人精细操作性能。
📝 Abstract
Vision-language-action (VLA) models provide a scalable path toward generalist robotic manipulation by integrating visual perception, language understanding, and continuous action control. However, we reveal a critical limitation of VLA architectures: the action expert has limited access to the 3D geometric and 2D semantic information available in VLM features. This accessibility gap weakens perceptual grounding and limits performance on fine-grained robotic manipulation. To address this issue, we propose V-Link, which explicitly recovers visual representations during the vision-language (VL) to action (A) feature transfer. Specifically, V-Link learns complementary Spatial and Semantic Query representations within the VLM and injects them into Action DiT through asymmetric pathways. Semantic Queries complement the original VLM image tokens, whereas Spatial Queries provide dedicated geometric conditioning for spatially grounded action generation. Across LIBERO, LIBERO-Plus, and RoboTwin 2.0, our V-Link improves the average success rate over base model GR00T N1.6 by +1.9%, +31.2%, and +18.8%, respectively. On the AGIBOT A3 Ultra, V-Link further achieves gains of +20% and +24% on two real-world humanoid tasks.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action Models
Geometric Information
Semantic Information
Perceptual Grounding
Innovation

Methods, ideas, or system contributions that make the work stand out.

V-Link
Spatial Queries
Semantic Queries
Vision-Language-Action Models
Feature Transfer
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Yehao Lu
Yehao Lu
Zhejiang University
Autonomous Driving3D ReconstructionSwarm Robot
J
Jiarui Yang
The Hong Kong University of Science and Technology (Guangzhou)
Y
Yuning Su
Simon Fraser University
Y
Yufeng Xie
AGIBOT
Y
Yu Zhong
AGIBOT
Yazhou Zhang
Yazhou Zhang
Associate Professor, Tianjin University
Sentiment AnalysisQuantum CognitionSarcasm DetectionHumor Analysis
Haiyu Lan
Haiyu Lan
University of Calgary
PerceptionLocalization and Mapping for Autonomous VehicleRobotics.
K
Kaixiang Lu
AGIBOT
P
Peiwen Lin
AGIBOT
C
Chuang Wang
AGIBOT
Zequn Qin
Zequn Qin
Zhejiang University
computer visiondeep learningmachine learning
E
Enyu Li
AGIBOT
X
Xi Li
Zhejiang University