H2R-Bench:世界模型中人類到機器人操作影片生成的基準測試
H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models
August 13, 2026
作者: Dingyi Rong, Yue Shi, Chaofan Ma, Jiezhang Cao, Zongrui Wang, Zeyu Zhang, Yao Mu, Guangtao Zhai, Ning Liu
cs.AI
摘要
大規模操作數據對於機器人學習至關重要,然而收集機器人示範數據仍然成本高昂且難以規模化。與此同時,豐富的第一人稱人類操作視頻提供了大量的行為經驗,但由於人類手部與機器人末端執行器之間的差異,跨具身遷移仍然充滿挑戰。視頻世界模型的最新進展為從人類觀測中合成以機器人為中心的操作視頻提供了有前景的途徑,但其跨具身遷移能力在很大程度上尚未被探索。為此,我們提出H2R-Bench,一個用於評估跨具身人機操作視頻生成的基準,其中模型需要在指定具身條件下將第一人稱人類示範轉換為機器人操作視頻。每個基準實例包含一段人類示範視頻、目標具身約束,以及涵蓋任務目標、動作事件、功能接觸和物體響應的源端對齊注釋。H2R-Bench從五個維度評估生成的視頻,包括目標狀態完成度、動作事件完成度、功能接觸遷移、具身正確性和一般視頻質量。我們對六個操作類別和兩種機器人具身條件下的十一個最先進視頻生成模型進行了基準評估。評估結果表明,當前的視頻世界模型在人機操作遷移方面仍然存在局限性:即使是領先模型也經常在具身一致性、功能交互和任務執行方面失敗。H2R-Bench提供了一個系統化的診斷框架,用於評估視頻世界模型能否彌合人機具身差距,並將人類操作觀測轉化為以機器人為中心的訓練資源。
English
Large-scale manipulation data is essential for robot learning, yet collecting robot demonstrations remains expensive and difficult to scale. Meanwhile, abundant egocentric human manipulation videos provide rich behavioral experiences, but transferring them across embodiments remains challenging due to differences between human hands and robotic end-effectors. Recent advances in video world models offer a promising pathway to synthesize robot-centric manipulation videos from human observations, while their cross-embodiment transfer capability remains largely unexplored. Therefore, we introduce H2R-Bench, a benchmark for evaluating cross-embodiment human-to-robot manipulation video generation, where models transform egocentric human demonstrations into robot manipulation videos under specified embodiments. Each benchmark instance contains a human demonstration video, target embodiment constraints, and source-grounded annotations covering task goals, action events, functional contacts, and object responses. H2R-Bench evaluates generated videos through five dimensions, including goal-state completion, action-event completion, functional contact transfer, embodiment correctness, and general video quality. We benchmark eleven state-of-the-art video generation models across six manipulation families and two robot embodiments. Our evaluation reveals that current video world models remain limited in human-to-robot manipulation transfer: even leading models often fail in embodiment consistency, functional interaction, and task execution. H2R-Bench provides a systematic diagnostic framework for evaluating whether video world models can bridge the human-to-robot embodiment gap and convert human manipulation observations into robot-centric training resources.