CIDER: Continual Interactive Distillation for Embodied Reinforcement Learning

📅 2026-08-22
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过CIDER框架解决持续学习中技能遗忘问题,采用交互式蒸馏和梯度路由方法保持旧技能同时学习新技能。
📝 Abstract
Human-in-the-loop real-world reinforcement learning enables rapid acquisition of effective robotic manipulation policies for individual tasks, often within tens of minutes. Yet it remains unclear how to extend this paradigm to continual learning, where a single policy must acquire new skills without losing previously learned behaviors. Existing real-world continual learning methods do not explicitly constrain prior behaviors, leading to severe catastrophic forgetting. We introduce Continual Interactive Distillation for Embodied Reinforcement Learning (CIDER), a continual reinforcement learning framework that freezes the accumulated historical policy as a teacher before learning each new task and interleaves task learning with distillation-based retention. We further introduce gradient routing to separate the gradients used for acquiring new tasks from those used for preserving prior behaviors. We evaluate our method with a single shared actor on six real-world household and industrial manipulation tasks. Interactive Distillation maintains high measured success on previously learned tasks across our six-task real-robot sequence while acquiring each new task in 10 to 20 minutes, whereas every baseline forgets at least one previous task. Additional ablations reveal the key design choices that govern the tradeoff between stability and plasticity in real-world continual reinforcement learning.
Problem

Research questions and friction points this paper is trying to address.

continual learning
catastrophic forgetting
reinforcement learning
interactive distillation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Continual Learning
Interactive Distillation
Gradient Routing
Catastrophic Forgetting
🔎 Similar Papers
2024-10-04International Conference on Learning RepresentationsCitations: 0
💼 Related Jobs
No related jobs found.
H
Houlin Li
AgiBot, Shanghai Jiao Tong University
M
Minghui Xu
AgiBot, Shanghai Jiao Tong University
G
Guo Xu
AgiBot
X
Xuan Du
AgiBot
X
Xiaohan Yan
AgiBot
Chun Wang
Chun Wang
University of Washington
Psychometricsmeasurementapplied statistics
Y
Yuxiang Yan
AgiBot
S
Shukai Yang
AgiBot
Y
Yongcheng Liu
AgiBot
W
Wei Shan
AgiBot
Maoqing Yao
Maoqing Yao
Google