Dynamics-Aware Meta-Imitation for Generalization to Unseen Robotic Manipulation

📅 2026-07-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limited generalization of robotic imitation learning in unseen tasks, which stems from data scarcity and environmental discrepancies. To overcome this, the authors propose the DAMI framework, which leverages meta-learning to construct a shared skill space and introduces three key components: a Visual-Motor Trajectory (VMT) module to model spatiotemporal dynamics, an Unpaired Unified Task (U2T) module to fuse multimodal observations without requiring paired demonstrations, and a Task-Conditioned Feature Modulation (TCFM) mechanism that emphasizes task-essential features over superficial cues. DAMI enables rapid adaptation to new tasks with only a few samples and no need for task-aligned demonstrations. Experimental results demonstrate that DAMI significantly outperforms existing methods in both simulation and real-world settings, achieving strong performance on seen tasks and exceptional generalization to unseen tasks after minimal fine-tuning.
📝 Abstract
Imitation Learning aims to learn skills from extensive observations and demonstrations for robots, so it suffers from data scarcity and environment generalization. The existing methods predominantly focus on imitation from in-domain tasks and consequently struggle with generalization to unseen tasks. To bridge this generalization gap, we propose the \textbf{D}ynamics-\textbf{A}ware \textbf{M}eta-\textbf{I}mitation (DAMI) framework. By integrating meta-learning to construct a shared skill space, DAMI equips agents for rapid adaptation to novel tasks. We introduce the Visual-Motor Trajectory (VMT) module to capture complex spatio-temporal dynamics within the task latent space. Furthermore, we propose the Unpaired Unified Task (U2T) block to fuse unstructured multimodal observations. To coordinate these representations, we integrate a Task-Conditioned Feature Modulation (TCFM) mechanism customized for modulating low-level 3D features. By capturing intrinsic dynamics from a random complete reference demonstration, our framework learns the underlying task logic rather than memorizing static cues, ensuring effective generalization. Extensive experiments in both simulation and real-world settings demonstrate that our approach outperforms state-of-the-art baselines regarding direct inference on seen tasks and adaptation to unseen tasks via few-shot fine-tuning.
Problem

Research questions and friction points this paper is trying to address.

Imitation Learning
Generalization
Robotic Manipulation
Meta-Learning
Unseen Tasks
Innovation

Methods, ideas, or system contributions that make the work stand out.

meta-imitation learning
dynamics-aware representation
visual-motor trajectory
task-conditioned feature modulation
unseen task generalization
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Zhenduo Shang
State Key Laboratory of Robotics and Intelligent Systems, Shenyang Institute of Automation, Chinese Academy of Sciences, Shenyang 110016, China; also with University of Chinese Academy of Sciences, Beijing 100049, China
X
Xiyao Liu
State Key Laboratory of Robotics and Intelligent Systems, Shenyang Institute of Automation, Chinese Academy of Sciences, Shenyang 110016, China
B
Bohan Li
Shenyang University of Technology
Xudong Wang
Xudong Wang
The Chinese University of Hong Kong, Shenzhen
Machine LearningGraph LearningSmart Grid
T
Teng Ren
Shenyang University of Technology
Lianqing Liu
Lianqing Liu
Professor, Shenyang Institute of Automation, Chinese Academy of Sciences
Biosyncretic RobotMicro/Nano RoboticsIntelligent Machine
Zhi Han
Zhi Han
SIA, CAS
Computer Vision