Dream2Reward: Transition-Alignment Reward Models from Positive Demonstrations for Robotic Manipulation

📅 2026-08-19
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出Dream2Reward方法,通过正面示例学习语言条件的成功潜在转换场,以解决机器人操作中需要密集奖励的问题,提供比基于进度的方法更有效的反馈。
📝 Abstract
Learning robotic policies requires dense rewards that remain informative when behavior departs from successful demonstrations. Progress-based rewards estimate how far an observation has advanced along a nominal successful trajectory, but may remain high after an incorrect transition. We introduce Dream2Reward, which learns a language-conditioned successful latent transition field from positive demonstrations. Given the visual history up to a transition start, the model predicts the latent displacement associated with successful execution and scores the observed displacement through signed directional and symmetric magnitude agreement. This transition-level comparison penalizes wrong-direction, overshooting, and stagnant motion even when the resulting observation appears to show progress. Dream2Reward requires no failure annotations, progress labels, or synthetic negatives, and produces a dense causal reward. Across mechanism diagnostics and shared-trajectory evaluations, it provides stronger success-failure separation and more informative feedback on low-quality behavior than progress-based alternatives. Across online and offline policy learning, the same frozen reward model reduces reward hacking and supports stronger downstream performance, including in real-robot manipulation. These results show that comparing realized motion with predicted successful change provides an effective way to convert positive demonstrations into dense rewards for robot learning.
Problem

Research questions and friction points this paper is trying to address.

robotic manipulation
dense rewards
positive demonstrations
transition-level comparison
reward hacking
Innovation

Methods, ideas, or system contributions that make the work stand out.

language-conditioned
latent transition field
positive demonstrations
dense causal reward
robotic manipulation
💼 Related Jobs
No related jobs found.
Haoyu Zhang
Haoyu Zhang
Institute of Automation, Chinese Academy of Sciences
Roboticscontrol theorydeep learningreinforcement learning
Z
Zecui Zeng
JD Explore Academy, China
B
Bin Wang
JD Explore Academy, China
L
Lusong Li
JD Explore Academy, China
Liang Lin
Liang Lin
Fellow of IEEE/IAPR, Professor of Computer Science, Sun Yat-sen University
Embodied AICausal Inference and LearningMultimodal Data Analysis
L
Long Cheng
School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing 100049, China; and State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, Beijing 100190, China