Robo-Dopamine 2.0: History-Conditioned and OOD-Aware Process Reward Modeling for Robotic Manipulation

📅 2026-08-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses error accumulation in robotic Vision-Language-Action (VLA) models and the limited discriminative power of existing reward models under out-of-distribution (OOD) conditions by proposing a history-conditioned, OOD-aware process reward model. We introduce novel mechanisms including history-conditioned pairwise prediction, symbolic progress space modeling, and signed-hop curriculum learning, alongside a dedicated benchmark dataset, to enhance policy robustness. Experimental results demonstrate that the proposed approach achieves a visual temporal consistency score of 0.9872, an 86.8% success rate on RobotWin, and an 88.75% success rate in real-world insertion tasks. These findings indicate significant improvements in both generalization capability and execution reliability for complex manipulation tasks, effectively mitigating compounding errors and enhancing adaptability in OOD scenarios.
📝 Abstract
Vision-language-action (VLA) models improve robotic manipulation but remain vulnerable to compounding errors, scene changes, and off-trajectory states. Reinforcement learning can refine pretrained VLA policies, yet sparse success signals hinder exploration, while engineered dense rewards are costly and task-specific. Existing learned visual reward models often rely on static before-after observations, causing temporal ambiguity and weak discrimination between robustness-preserving variations and task-invalid failures under out-of-distribution (OOD) execution. We introduce Robo-Dopamine 2.0, a history- and OOD-aware process reward model with a pairwise prediction interface. It combines (1) history-conditioned pairwise rewards that use source-aligned reference panels for synthetic OOD queries and observed rollout history for online queries, while preserving the queried endpoints, and (2) an OOD-aware signed progress space that represents valid progress, robustness, failure, and recovery. A Signed-Hop Curriculum with transition-aware replay learns coarse execution ordering before fine-grained progress calibration. We also construct an OOD trajectory dataset and a five-family benchmark. Reference panels improve mean visual order consistency (VOC) from 0.967 to 0.986 and OOD-robust VOC from 0.906 to 0.958. With the same 400K pairwise-reward budget, Signed-Hop training with 25% replay reaches 0.9872 mean VOC, compared with 0.9858 for a matched-pool shuffled control. In downstream reinforcement learning, the full model achieves 86.8% mean RoboTwin success and 71/80 successful real-world insertions.
Problem

Research questions and friction points this paper is trying to address.

Process Reward Modeling
Robotic Manipulation
Out-of-Distribution
Vision-Language-Action Models
Reinforcement Learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Process Reward Modeling
Out-of-Distribution Awareness
History-Conditioned Pairwise Rewards
Signed Progress Space
Signed-Hop Curriculum
💼 Related Jobs
No related jobs found.
Yijie Xu
Yijie Xu
Hong Kong University of Science and Technology (Guangzhou)
Data MiningNatural Language ProcessingLarge Language Models
H
Haopeng Jin
Beijing University of Posts and Telecommunications
R
Run Zhou
Renmin University of China
S
Shengbang Liu
Sun Yat-sen University
Sixiang Chen
Sixiang Chen
The Hong Kong University of Science and Technology (Guangzhou)
Computer VisionImage RestorationAIGCMLLM
H
Hongyang Cheng
EvoPhys AI
S
Sicheng Hu
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University
P
Peterson Co
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University
J
Jinwen Luo
Tencent
Huajie Tan
Huajie Tan
Peking University
Embodied AIFoundation Models
Shanghang Zhang
Shanghang Zhang
Peking University
Embodied AIFoundation Models