PRM-as-a-Judge 1.5: A Toolkit for Robot Process Assessment
Yuyang Liu, Yanqing Shen, Ruike Chen, Jifan Zhao, Yuxuan Tian, Yichi Zhang, Tianfeng Long, Zixuan Yin, Yipu Wang, Ziheng Qin, Wenxing Tan, Yang Shi, Mingyu Cao, Runze Xiao, Ziqi Wang, Zhixin Yin, Shiwei Chu, Yi-Fan Zhang, Yao Mu, Yuheng Ji, Yihao Wang, Jun Yan, Zhongyuan Wang, Pengwei Wang, Xiaolong Zheng
arXiv:2608.14284 · 2026-08-17 공개 · arXiv · PDF
benchmarking reproducible-evaluation process-reward-models visualization-tools failure-recovery reliability-evaluation robot-process-assessment embodied-models
Abstract
Fine-grained robotic evaluation matters for understanding embodied models, going beyond binary success rates and rule-based process scores. We present PRM-as-a-Judge 1.5, a toolkit for robot process assessment that turns rollout videos into dense progress curves and derives multiple fine metrics. PRM-as-a-Judge 1.5 introduces three metrics, building on version 1.0, that characterize failure-side progress, post-drawdown recovery, and success-side execution quality, helping users understand embodied model capability. Based on the rollout videos from benchmarks, we perform a comprehensive assessment of the embodied models, providing some fine-grained metric results and key findings. We further introduce RoboPulse++ to evaluate the reliability of process reward models (PRM), providing evaluators with a more accurate testing platform. Moreover, we release a user-friendly assessment suite, including the benchmark, metric implementation, and visualization tools, to support reproducible manipulation process evaluation. We call on the community to rethink how robots are evaluated and establish transparent, procedural, and reproducible assessment as a foundation for the next generation of embodied intelligence.
한국어 요약
한 줄 요약
PRM-as-a-Judge 1.5는 로봇 실행 과정을 정밀하게 평가하는 툴킷으로, 실패 측 진행, 회복, 성공 측 실행 품질을 측정한다.
핵심 기여도
- PRM-as-a-Judge 1.5는 1.0 버전에 이어 실패 측 진행, 회복, 성공 측 실행 품질을 측정하는 세 가지 새로운 메트릭을 추가.
- RoboPulse++를 도입하여 PRM의 신뢰도를 평가하는 테스트 플랫폼 제공.
- 벤치마크, 메트릭 구현, 시각화 도구를 포함한 오픈소스 평가 툴킷을 공개.
- 주요 로봇 조작 벤치마크에서 실행 과정을 분석하여 세부 실행 품질을 평가.
핵심 아이디어
기존 로봇 평가 방식은 이진 성공률이나 수동 규칙 기반 점수에 의존하여 실행 과정의 세부 정보를 무시한다. PRM-as-a-Judge 1.5는 실행 영상에서 밀집된 진행 곡선을 생성하고, 이를 기반으로 세 가지 조건 메트릭을 도출하여 로봇의 실행 품질을 정밀하게 평가한다. 이는 실패 시에도 진행 수준을 분석하고, 성공 시에도 실행 효율성을 평가할 수 있게 한다. OPD (Outcome–Process–Diagnosis) 메트릭 시스템은 실패 측 진행, 회복, 성공 측 품질을 구분하여 모델의 능력을 다각도로 해석한다.
기술적 접근법
- **PRM (Process Reward Model)**: 실행 영상에서 밀집된 진행 곡선을 생성.
- **OPD 메트릭 시스템**: Outcome (결과), Process (과정), Diagnosis (진단)을 기반으로 세부 평가.
- **RoboPulse++**: PRM의 신뢰도를 평가하는 인터벌 기반 벤치마크.
- **RoboDopamine (Forward)**: 기본 평가 모델로 사용.
- **시각화 도구**: 실행 과정을 시각적으로 분석할 수 있도록 지원.
- **재현 가능한 평가**: 벤치마크, 메트릭, 도구를 오픈소스로 공개하여 재현 가능성을 확보.
주요 결과
- **실패 측 진행**: 실패한 실행에서도 진행 수준을 측정하여 모델의 능력을 분석.
- **회복 행동**: 실행 중 회복 능력을 평가하여 안정성을 측정.
- **성공 측 품질**: 성공한 실행에서도 효율성, 정밀도 등을 평가.
- **RoboPulse++**: PRM의 신뢰도를 평가하여 1.6× 가속된 평가 성능을 보여.
- **실행 곡선 분석**: 주요 벤치마크에서 실행 곡선을 통해 실패 패턴을 식별.
의의 및 한계
PRM-as-a-Judge 1.5는 로봇 평가를 이진 성공률에서 벗어나 실행 과정을 정밀하게 분석할 수 있는 기반을 제공한다. 특히 실패 측 진행과 회복 능력을 평가함으로써 모델의 안정성과 유연성을 측정할 수 있다. 그러나 현재는 실행 영상 기반 평가에 의존하며, 실시간 실행 환경에서는 제한이 있을 수 있다. 또한, PRM의 정확도는 훈련 데이터의 질에 크게 의존하므로, 더 다양한 실패 및 회복 데이터가 필요하다.
실용적 활용
PRM-as-a-Judge 1.5는 로봇 제어, 자율 시스템 개발, 로봇 학습 연구 등에서 실행 과정을 정밀히 평가하는 데 활용될 수 있다. 특히 실패 분석과 회복 능력을 평가하여 안정적인 로봇 시스템 설계에 기여할 수 있다. 오픈소스 툴킷은 연구자와 엔지니어가 재현 가능한 평가를 수행할 수 있도록 지원한다.