GIFT: Guided Intermediate Feature Training via Action-Oriented Structural Supervision for Robotic Manipulation

📅 2026-09-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过GIFT框架解决视觉丰富性和控制实用性之间的不匹配问题,利用几何对齐、可承受性预测和目标区域重建方法提升机器人操作性能。
📝 Abstract
Vision-language pre-training and predictive world modeling provide robot policies with rich semantic and dynamic visual features, but their native action and visual-prediction objectives may omit critical physical and task structure while retaining control-irrelevant visual redundancy. We call this mismatch between visual richness and control utility the action-sufficiency gap. We investigate whether this gap can be bridged by guiding intermediate features to preserve three control-relevant structure in robotic manipulation: geometry governing motion feasibility, affordance encoding instruction-relevant entities, and goals grounding instructions in task-relevant regions. To this end, we present GIFT (Guided Intermediate Feature Training), an architecture-flexible framework for learning intermediate features that translates these structures into training-time constraints through geometry alignment, affordance prediction, and goal-region reconstruction. We instantiate GIFT in a Vision-Language-Action (VLA) policy, a direct-action World-Action Model (WAM), and an inverse-dynamics WAM while retaining each model's action formulation. Under zero-shot transfer to LIBERO-Plus, GIFT-VLA, GIFT-WAM-Fast, and GIFT-WAM-IDM outperform StarVLA-OFT, Fast-WAM, and Fast-WAM-IDM by 4.6, 12.6, and 5.2 points, reaching 79.6%, 72.6%, and 87.8%, respectively. On RoboCasa, the three GIFT variants reach 61.4%, 83.6%, and 82.3%, outperforming their counterparts by 12.6, 9.0, and 8.4 points, respectively. Together, these results establish learning functionally structured intermediate features as a reusable principle across model-specific action formulations, with especially large gains on articulated-object tasks and high-precision real-world manipulation under unseen visual and spatial perturbations. Project page: https://openphoenix-team.github.io/GIFT-pages.
Problem

Research questions and friction points this paper is trying to address.

robotic manipulation
action-sufficiency gap
visual redundancy
control-irrelevant
task structure
Innovation

Methods, ideas, or system contributions that make the work stand out.

GIFT
action-sufficiency gap
geometry alignment
affordance prediction
goal-region reconstruction
💼 Related Jobs
No related jobs found.
Yupeng Zheng
Yupeng Zheng
Institute of Automation, Chinese Academy of Sciences
X
Xiang Li
Tsinghua University
Songen Gu
Songen Gu
UCAS
Robotics3D Vision
Yuhang Zheng
Yuhang Zheng
NUS; TARS AI
Robotics3D Vision
S
Shuai Tian
Institute of Automation, Chinese Academy of Sciences; University of Chinese Academy of Sciences
W
Weize Li
National University of Singapore
Linbo Wang
Linbo Wang
University of Toronto
C
Chaoyue Li
Institute of Automation, Chinese Academy of Sciences; University of Chinese Academy of Sciences
Qichao Zhang
Qichao Zhang
中国科学院自动化研究所
人工智能 强化学习 博弈论 自适应动态规划
Haoran Li
Haoran Li
Institute of Automation,Chinese Academy of Sciences
Artificial IntelligenceRoboticsReinforcement LearningEmbodied Intelligence
Z
Zhongpu Xia
Institute of Automation, Chinese Academy of Sciences
Y
Ya-Qin Zhang
Tsinghua University
S
Shuicheng Yan
National University of Singapore
Dongbin Zhao
Dongbin Zhao
Institute of Automation, Chinese Academy of Sciences
Deep Reinforcement LearningAdaptive Dynamic ProgrammingGame AISmart drivingrobotics