TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models

πŸ“… 2026-08-09
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
Existing robotic reward models struggle to simultaneously maintain pointwise scoring accuracy and pairwise preference consistency in long-horizon tasks, leading to training noise and performance degradation. This work proposes Preference-Ordered Isotonic Score Editing (POISE), a novel method that achieves conflict-free alignment between these two signal types for the first time, effectively resolving the score-preference reversal problem. Evaluated on a unified four-paradigm dataset and leveraging a vision-language model with video question-answering supervision and the TrustJudge reasoning aggregation strategy, Qwen3-VL-4B calibrated by POISE attains a reward accuracy of 77.96% and improves score-preference consistency to 71.90%. Further integration of TrustJudge elevates the overall performance to 78.57%, surpassing the teacher model.
πŸ“ Abstract
Reward models are a bottleneck for reinforcement learning in embodied AI. Long-horizon robotic manipulation requires scalable vision feedback beyond handcrafted rewards or task-specific annotations. Existing open-source VLM reward judges like RoboReward adopt simple 1--5 trajectory progress scoring, lacking pairwise preferences for RLHF, DPO and Bradley-Terry frameworks, while failing to optimize video scene understanding. Augmenting RoboReward with pairwise comparison and video-QA supervision causes inconsistency between pairwise preferences and pointwise scores, introducing training noise and hurting downstream performance---an issue aggregation methods such as TrustJudge cannot resolve. To address this, we propose TrustRoboReward, a multi-paradigm reward modeling framework equipped with Preference-Ordered Isotonic Score Editing (POISE). We construct a unified four-paradigm dataset with trajectory progress scoring (Score-A), video-QA answer quality scoring (Score-B), and their pairwise counterparts (Pair-A, Pair-B). Pairwise labels align better with human judgment than pointwise scores, inspiring us to calibrate pointwise scores to avoid score-pair reversals against pairwise preferences. POISE rectifies pointwise scores and eliminates cross-paradigm reversal conflicts unresolved by TrustJudge. Theoretically, POISE reduces score-pair reversal conflicts from 20.15% to 0%, whereas TrustJudge retains 20.46% conflicts on the same corpus. Evaluated on our benchmark, Qwen3-VL-4B trained with POISE achieves an overall reward score of 77.96%, nearly matching GPT-5-mini (78.09%, gap 0.13%) and outperforming the strongest RoboReward-4B baseline by 10.13%. It also lifts test-time score-pair consistency to 71.90%, exceeding RoboReward-4B (57.26%) and GPT-5-mini (68.09%). Integrating TrustJudge aggregation during inference boosts the overall score to 78.57%, surpassing the GPT-5-mini teacher model.
Problem

Research questions and friction points this paper is trying to address.

reward models
pairwise preferences
pointwise scores
score-pair inconsistency
robotic reinforcement learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Preference-Ordered Isotonic Score Editing
multi-paradigm reward modeling
score-pair consistency
robotic reward models
vision-language models
Y
Yidong Wang
Peking University
Y
Yan Zhan
Peking University
Z
Ziteng Feng
University of Science and Technology of China
Zhenyu Cui
Zhenyu Cui
Associate Professor, School of Business, Stevens Institute of Technology
Financial EngineeringFinancial TechnologyDerivative pricingInsurance Analytics
Z
Ziyi Zhou
Southern University of Science and Technology
R
Renzhao Liang
Beijing University of Aeronautics and Astronautics
J
Jiaxuan Zhu
Southeast University
Z
Zilei Yang
Beijing University of Aeronautics and Astronautics
Yiran Zhao
Yiran Zhao
National University of Singapore
ReasoningEfficiencyMultilingualAlignment
Z
Zhongkuan Mao
Sichuan University
B
Bo Jia
Beijing University of Posts and Telecommunications
H
Hanchu Ni
Peking University
C
Chenggang Xie
Beijing University of Aeronautics and Astronautics
B
Biao Liu
Southeast University
Yi Zhang
Yi Zhang
Beijing Institute of Technology
Y
Yong Dai
Beijing Innovation Center of Humanoid Robotics
X
Xiaozhu Ju
Beijing Innovation Center of Humanoid Robotics
Wei Ye
Wei Ye
Peking University
Software EngineeringNatural Language Processing
Shikun Zhang
Shikun Zhang
εŒ—δΊ¬ε€§ε­¦