PRM-as-a-Judge 1.5: A Toolkit for Robot Process Assessment

📅 2026-08-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of embodied model evaluation that relies on binary success rates and lacks fine-grained process analysis by proposing a video-based dense progress curve generation method. Introducing three novel metrics—progress on failure, recovery from backtracking, and execution quality—the authors construct the RoboPulse++ benchmark platform alongside a visual evaluation toolkit. This work facilitates a paradigm shift from outcome-oriented to process-aware assessment, revealing intricate behavioral details of embodied models. Furthermore, the complete toolset is open-sourced to support transparent and reproducible fine-grained manipulation evaluation, thereby advancing the standardization of embodied intelligence benchmarks.
📝 Abstract
Fine-grained robotic evaluation matters for understanding embodied models, going beyond binary success rates and rule-based process scores. We present PRM-as-a-Judge 1.5, a toolkit for robot process assessment that turns rollout videos into dense progress curves and derives multiple fine metrics. PRM-as-a-Judge 1.5 introduces three metrics, building on version 1.0, that characterize failure-side progress, post-drawdown recovery, and success-side execution quality, helping users understand embodied model capability. Based on the rollout videos from benchmarks, we perform a comprehensive assessment of the embodied models, providing some fine-grained metric results and key findings. We further introduce RoboPulse++ to evaluate the reliability of process reward models (PRM), providing evaluators with a more accurate testing platform. Moreover, we release a user-friendly assessment suite, including the benchmark, metric implementation, and visualization tools, to support reproducible manipulation process evaluation. We call on the community to rethink how robots are evaluated and establish transparent, procedural, and reproducible assessment as a foundation for the next generation of embodied intelligence.
Problem

Research questions and friction points this paper is trying to address.

Fine-grained robotic evaluation
Embodied models
Process reward models
Manipulation process assessment
Reproducible evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

PRM-as-a-Judge
Fine-grained Evaluation
Process Reward Model
RoboPulse++
Dense Progress Curves
Y
Yuyang Liu
Yanqing Shen
Yanqing Shen
Xi'an jiaotong University, ETH
visual place recognitionimage representation
R
Ruike Chen
J
Jifan Zhao
Yuxuan Tian
Yuxuan Tian
Peking University
Y
Yichi Zhang
T
Tianfeng Long
Z
Zixuan Yin
Y
Yipu Wang
Z
Ziheng Qin
W
Wenxing Tan
Y
Yang Shi
M
Mingyu Cao
R
Runze Xiao
Z
Ziqi Wang
Z
Zhixin Yin
S
Shiwei Chu
Yi-Fan Zhang
Yi-Fan Zhang
Institute of Automation, Chinese Academy of Sciences
Computer VisionMultimodalityAlignmentMachine Learning
Y
Yao Mu
Yuheng Ji
Yuheng Ji
Institute of Automation, Chinese Academy of Sciences
Embodied AIComputer Vision
Y
Yihao Wang
J
Jun Yan
Z
Zhongyuan Wang
Pengwei Wang
Pengwei Wang
University of Calgary
Computer Science Security
X
Xiaolong Zheng