ICI-VLA: In-Context Imitation with Spatiotemporally Aligned Demonstrations for Vision-Language-Action Models

📅 2026-09-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出ICI-VLA框架,通过在测试时使用时空对齐的演示来解决视觉-语言-动作模型快速适应新任务的问题,无需额外梯度更新。
📝 Abstract
Vision-Language-Action (VLA) policies are commonly adapted to new manipulation settings through additional gradient updates, which limits rapid deployment when task-specific data or compute is scarce. We present ICI-VLA, a training and retrieval framework that equips a text-action VLM with few-shot test-time adaptation through in-context demonstrations. Unlike mainstream VLA designs based on action-specific multimodal fusion, ICI-VLA retains the native text-generation interface. ICI-VLA updates its parameters only during offline training; at inference, the policy remains fixed and conditions action generation on retrieved micro-demonstrations. The framework decomposes long trajectories into short, semantically labeled examples and trains an RD-Encoder with positives mined by Dynamic Time Warping (DTW), aligning the retrieved context with the phase and geometry of the current subtask. We further introduce Target Action Masking, a context-corruption objective designed to reduce direct action copying and increase reliance on the current observation. ICI-VLA reaches average success rates of 97.7% on LIBERO and 60.4% on RoboTwin 2.0, exceeding the highest reported baseline average on RoboTwin 2.0 by 19.3 percentage points. It also achieves 83.2% across four physical tasks. These results indicate that a fixed VLA policy can benefit from conditioning on spatiotemporally aligned demonstrations at test time.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action
test-time adaptation
few-shot learning
spatiotemporally aligned demonstrations
Innovation

Methods, ideas, or system contributions that make the work stand out.

In-Context Imitation
Spatiotemporally Aligned Demonstrations
Target Action Masking
Dynamic Time Warping (DTW)
S
Songhua Yang
Wuhan University, Wuhan, China
Z
Ziyu Liu
Wuhan University, Wuhan, China
X
Xuetao Li
Wuhan University, Wuhan, China
R
Ruqi Xiao
Wuhan University, Wuhan, China
K
Kangxin Zhu
Wuhan University, Wuhan, China
Miao Li
Miao Li
Professor, Wuhan University
RoboticsGraspingDexterous ManipulationLearning from Demonstration