DIA: Denoising Intermediate Advantage for Diffusion Policy Optimization

📅 2026-09-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出DIA方法,通过学习部分去噪动作的价值函数并结合环境级PPO优势,改进了扩散策略的微调,从而提高机器人操作性能。
📝 Abstract
Diffusion-based robot policies have become widely used in robotic manipulation, where they are typically trained with behavior cloning. However, policies trained purely from demonstrations are limited by the quality and coverage of the available data. Reinforcement learning can further improve the performance of these pretrained policies through interaction. A common approach is to use policy-gradient methods that formulate diffusion-policy fine-tuning as an outer environment MDP together with an inner denoising MDP. However, existing methods typically assign the same environment-level credit to all denoising steps used to construct an action chunk, without distinguishing which intermediate decisions contributed most to the final return. We introduce Denoising Intermediate Advantage (DIA), a policy-gradient method that learns a value function over partially denoised actions and uses it to construct a denoising- level advantage for each step of the generative process. DIA com- bines this inner credit signal with the standard environment-level PPO advantage, providing state-dependent credit throughout the denoising chain. Across Robomimic, FurnitureBench, Franka Kitchen, and D3IL, DIA consistently improves final performance over existing diffusion-policy fine-tuning methods. Beyond final reward, DIA reaches successful states more efficiently and can shift farther from the pretrained behavior distribution, enabling it to discover more effective and efficient task-level strategies and subtask sequences that baseline methods fail to reach.
Problem

Research questions and friction points this paper is trying to address.

diffusion-policy
reinforcement learning
behavior cloning
credit assignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Denoising Intermediate Advantage
Policy Gradient
Diffusion Policies
Reinforcement Learning
Behavior Cloning
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Arjun Sohal
University of Toronto, Toronto, ON, Canada
Y
Yuchi Zhao
University of Toronto, Toronto, ON, Canada; Vector Institute for Artificial Intelligence, Toronto, ON, Canada
Miroslav Bogdanovic
Miroslav Bogdanovic
University of Toronto
Reinforcement LearningDeep LearningRobotics
A
Alán Aspuru-Guzik
University of Toronto, Toronto, ON, Canada; Vector Institute for Artificial Intelligence, Toronto, ON, Canada; Acceleration Consortium, University of Toronto, Toronto, ON, Canada; Canadian Institute for Advanced Research (CIFAR), Toronto, ON, Canada; NVIDIA, Toronto, ON, Canada