H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models

Dingyi Rong, Yue Shi, Chaofan Ma, Jiezhang Cao, Zongrui Wang, Zeyu Zhang, Yao Mu, Guangtao Zhai, Ning Liu

arXiv:2608.13049 · 2026-08-14 공개 · arXiv · PDF

video-generation world-models robot-manipulation robot-learning video-quality cross-embodiment task-execution manipulation

Abstract

Large-scale manipulation data is essential for robot learning, yet collecting robot demonstrations remains expensive and difficult to scale. Meanwhile, abundant egocentric human manipulation videos provide rich behavioral experiences, but transferring them across embodiments remains challenging due to differences between human hands and robotic end-effectors. Recent advances in video world models offer a promising pathway to synthesize robot-centric manipulation videos from human observations, while their cross-embodiment transfer capability remains largely unexplored. Therefore, we introduce H2R-Bench, a benchmark for evaluating cross-embodiment human-to-robot manipulation video generation, where models transform egocentric human demonstrations into robot manipulation videos under specified embodiments. Each benchmark instance contains a human demonstration video, target embodiment constraints, and source-grounded annotations covering task goals, action events, functional contacts, and object responses. H2R-Bench evaluates generated videos through five dimensions, including goal-state completion, action-event completion, functional contact transfer, embodiment correctness, and general video quality. We benchmark eleven state-of-the-art video generation models across six manipulation families and two robot embodiments. Our evaluation reveals that current video world models remain limited in human-to-robot manipulation transfer: even leading models often fail in embodiment consistency, functional interaction, and task execution. H2R-Bench provides a systematic diagnostic framework for evaluating whether video world models can bridge the human-to-robot embodiment gap and convert human manipulation observations into robot-centric training resources.

한국어 요약

한 줄 요약

H2R-Bench는 인간 조작 영상을 로봇 중심 조작 영상으로 전환하는 능력을 평가하는 벤치마크로, 6개 조작 유형과 2개 로봇 구조에서 11개 모델을 평가한다.

핵심 기여도

핵심 아이디어

H2R-Bench는 인간 조작 영상에서 로봇 조작 영상으로의 전환을 평가하기 위해 설계된 벤치마크로, 기존의 텍스트 또는 로봇 중심 조건에 의존하는 평가와 달리, 특정 인간 영상에 기반한 전환을 요구한다. 이는 로봇이 인간의 조작 의도, 기능적 접촉, 객체 상태 변화를 정확히 반영하면서도, 로봇의 구조와 종단기구(end-effector)에 맞게 실행해야 한다는 점에서 차별화된다. 핵심 통찰은, 시각적 품질(M5)이 높아도 실제 로봇 전환 능력(H2RCore)은 낮을 수 있으며, 이는 기존 영상 생성 모델이 인간-로봇 체제 간의 차이를 충분히 이해하지 못함을 시사한다.

기술적 접근법

주요 결과

의의 및 한계

H2R-Bench는 인간 조작 영상에서 로봇 조작 영상으로의 전환 능력을 체계적으로 평가하는 첫 번째 벤치마크로, 로봇 학습에서 데이터 부족 문제를 해결할 가능성을 제시한다. 특히, 기존 영상 생성 모델이 시각적 품질에 치우쳐 인간-로봇 체제 간의 차이를 충분히 반영하지 못함을 밝혀내며, 이 분야의 연구 방향을 재정립하는 데 기여한다. 그러나 H2R-Bench는 2개 로봇 구조만 포함하며, 더 다양한 로봇 형태나 환경에 대한 평가가 필요한 한계가 있다.

실용적 활용

H2R-Bench는 로봇 학습 데이터 생성, 인간-로봇 전이 학습, 로봇 행동 시뮬레이션 등에 활용될 수 있다. 특히, 대규모 인간 조작 영상을 로봇 중심 학습 자료로 전환하여, 로봇 학습의 효율성과 확장성을 높이는 데 기여할 수 있다.