H2R-Bench: 세계 모델에서의 인간-로봇 조작 비디오 생성 벤치마킹
H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models
August 13, 2026
저자: Dingyi Rong, Yue Shi, Chaofan Ma, Jiezhang Cao, Zongrui Wang, Zeyu Zhang, Yao Mu, Guangtao Zhai, Ning Liu
cs.AI
초록
대규모 조작 데이터는 로봇 학습에 필수적이지만, 로봇 시연 데이터를 수집하는 것은 여전히 비용이 많이 들고 확장이 어렵다. 한편, 풍부한 1인칭 인간 조작 비디오는 다양한 행동 경험을 제공하지만, 인간의 손과 로봇 엔드 이펙터 간의 차이로 인해 이를 다른 엠보디먼트로 전이하는 것은 여전히 어려운 과제로 남아 있다. 최근 비디오 월드 모델의 발전은 인간의 관찰로부터 로봇 중심의 조작 비디오를 합성할 수 있는 유망한 경로를 제시하지만, 이들의 교차 엠보디먼트 전이 능력은 아직 충분히 탐구되지 않았다. 이에 우리는 모델이 1인칭 인간 시연을 지정된 엠보디먼트 조건에서 로봇 조작 비디오로 변환하는 교차 엠보디먼트 인간-로봇 조작 비디오 생성을 평가하기 위한 벤치마크인 H2R-Bench를 제안한다. 각 벤치마크 인스턴스는 인간 시연 비디오, 대상 엠보디먼트 제약 조건, 그리고 작업 목표, 행동 이벤트, 기능적 접촉, 객체 반응을 포괄하는 소스 기반 주석으로 구성된다. H2R-Bench는 목표 상태 달성, 행동 이벤트 달성, 기능적 접촉 전이, 엠보디먼트 정확성, 일반 비디오 품질의 다섯 가지 차원을 통해 생성된 비디오를 평가한다. 우리는 여섯 가지 조작 유형과 두 가지 로봇 엠보디먼트에 걸쳐 열한 개의 최신 비디오 생성 모델을 평가한다. 평가 결과, 현재의 비디오 월드 모델은 인간-로봇 조작 전이에 있어 여전히 한계가 있음이 드러났다. 선도적인 모델조차도 엠보디먼트 일관성, 기능적 상호작용, 작업 실행에서 실패하는 경우가 많다. H2R-Bench는 비디오 월드 모델이 인간-로봇 간 엠보디먼트 격차를 해소하고 인간 조작 관찰을 로봇 중심의 훈련 자원으로 변환할 수 있는지 평가하기 위한 체계적인 진단 프레임워크를 제공한다.
English
Large-scale manipulation data is essential for robot learning, yet collecting robot demonstrations remains expensive and difficult to scale. Meanwhile, abundant egocentric human manipulation videos provide rich behavioral experiences, but transferring them across embodiments remains challenging due to differences between human hands and robotic end-effectors. Recent advances in video world models offer a promising pathway to synthesize robot-centric manipulation videos from human observations, while their cross-embodiment transfer capability remains largely unexplored. Therefore, we introduce H2R-Bench, a benchmark for evaluating cross-embodiment human-to-robot manipulation video generation, where models transform egocentric human demonstrations into robot manipulation videos under specified embodiments. Each benchmark instance contains a human demonstration video, target embodiment constraints, and source-grounded annotations covering task goals, action events, functional contacts, and object responses. H2R-Bench evaluates generated videos through five dimensions, including goal-state completion, action-event completion, functional contact transfer, embodiment correctness, and general video quality. We benchmark eleven state-of-the-art video generation models across six manipulation families and two robot embodiments. Our evaluation reveals that current video world models remain limited in human-to-robot manipulation transfer: even leading models often fail in embodiment consistency, functional interaction, and task execution. H2R-Bench provides a systematic diagnostic framework for evaluating whether video world models can bridge the human-to-robot embodiment gap and convert human manipulation observations into robot-centric training resources.