TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models
Existing robotic reward models struggle to simultaneously maintain pointwise scoring accuracy and pairwise preference consistency in long-horizon tasks, leading to training noise and performance degradation. This work proposes Preference-Ordered Isotonic Score Editing (POISE), a novel method that achieves conflict-free alignment between these two signal types for the first time, effectively resolving the score-preference reversal problem. Evaluated on a unified four-paradigm dataset and leveraging a vision-language model with video question-answering supervision and the TrustJudge reasoning aggregation strategy, Qwen3-VL-4B calibrated by POISE attains a reward accuracy of 77.96% and improves score-preference consistency to 71.90%. Further integration of TrustJudge elevates the overall performance to 78.57%, surpassing the teacher model.