H2R-Bench: ワールドモデルにおける人間からロボットへの操作ビデオ生成のベンチマーキング
H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models
August 13, 2026
著者: Dingyi Rong, Yue Shi, Chaofan Ma, Jiezhang Cao, Zongrui Wang, Zeyu Zhang, Yao Mu, Guangtao Zhai, Ning Liu
cs.AI
要旨
大規模な操作データはロボット学習に不可欠であるが、ロボットのデモンストレーション収集は依然として高コストであり、大規模化が困難である。一方、豊富に存在するエゴセントリックな人間の操作ビデオは豊かな行動経験を提供するものの、人間の手とロボットのエンドエフェクタの違いにより、それらをエンボディメント間で転移させることは依然として困難である。近年のビデオワールドモデルの進展は、人間の観察からロボット中心の操作ビデオを合成する有望な道筋を提供しているが、そのクロスエンボディメント転移能力はほとんど調査されていない。そこで我々は、エゴセントリックな人間のデモンストレーションを指定されたエンボディメント条件下でロボット操作ビデオに変換する、クロスエンボディメントの人間からロボットへの操作ビデオ生成を評価するためのベンチマークであるH2R-Benchを導入する。各ベンチマークインスタンスには、人間のデモンストレーションビデオ、ターゲットエンボディメントの制約、およびタスク目標、アクションイベント、機能的接触、物体応答をカバーするソースに基づくアノテーションが含まれる。H2R-Benchは、生成されたビデオを、ゴール状態の達成度、アクションイベントの達成度、機能的接触の転移、エンボディメントの正確性、および一般的なビデオ品質の5つの次元で評価する。我々は、6つの操作ファミリーと2つのロボットエンボディメントにわたって、11の最先端ビデオ生成モデルのベンチマーク評価を行う。評価の結果、現在のビデオワールドモデルは人間からロボットへの操作転移において依然として限界があることが明らかになった。すなわち、主要なモデルでさえ、エンボディメントの一貫性、機能的インタラクション、およびタスク実行において失敗することが多い。H2R-Benchは、ビデオワールドモデルが人間とロボットのエンボディメントギャップを橋渡しし、人間の操作観察をロボット中心のトレーニングリソースに変換できるかを評価するための体系的な診断フレームワークを提供する。
English
Large-scale manipulation data is essential for robot learning, yet collecting robot demonstrations remains expensive and difficult to scale. Meanwhile, abundant egocentric human manipulation videos provide rich behavioral experiences, but transferring them across embodiments remains challenging due to differences between human hands and robotic end-effectors. Recent advances in video world models offer a promising pathway to synthesize robot-centric manipulation videos from human observations, while their cross-embodiment transfer capability remains largely unexplored. Therefore, we introduce H2R-Bench, a benchmark for evaluating cross-embodiment human-to-robot manipulation video generation, where models transform egocentric human demonstrations into robot manipulation videos under specified embodiments. Each benchmark instance contains a human demonstration video, target embodiment constraints, and source-grounded annotations covering task goals, action events, functional contacts, and object responses. H2R-Bench evaluates generated videos through five dimensions, including goal-state completion, action-event completion, functional contact transfer, embodiment correctness, and general video quality. We benchmark eleven state-of-the-art video generation models across six manipulation families and two robot embodiments. Our evaluation reveals that current video world models remain limited in human-to-robot manipulation transfer: even leading models often fail in embodiment consistency, functional interaction, and task execution. H2R-Bench provides a systematic diagnostic framework for evaluating whether video world models can bridge the human-to-robot embodiment gap and convert human manipulation observations into robot-centric training resources.