H2R-Bench:世界模型中的人到机器人操作视频生成基准测试
H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models
August 13, 2026
作者: Dingyi Rong, Yue Shi, Chaofan Ma, Jiezhang Cao, Zongrui Wang, Zeyu Zhang, Yao Mu, Guangtao Zhai, Ning Liu
cs.AI
摘要
大规模操作数据对于机器人学习至关重要,然而采集机器人演示数据仍然成本高昂且难以规模化扩展。与此同时,丰富的第一人称人类操作视频提供了大量的行为经验,但由于人类手部与机器人末端执行器之间的差异,跨具身迁移仍然具有挑战性。视频世界模型的最新进展为实现从人类观察中合成以机器人为中心的操作视频提供了有前景的途径,但其跨具身迁移能力在很大程度上仍未得到探索。为此,我们提出了H2R-Bench,一个用于评估跨具身人至机器人操作视频生成的基准测试,其中模型在指定具身约束下将第一人称人类演示视频转化为机器人操作视频。该基准的每个实例包含一段人类演示视频、目标具身约束以及涵盖任务目标、动作事件、功能接触和物体响应的源标注信息。H2R-Bench从五个维度评估生成的视频,包括目标状态完成度、动作事件完成度、功能接触迁移、具身正确性和通用视频质量。我们评估了十一个最先进的视频生成模型,涵盖六个操作类别和两种机器人具身。评估结果表明,当前的视频世界模型在人至机器人操作迁移方面仍存在较大局限:即使是领先模型也常在具身一致性、功能交互和任务执行方面表现不足。H2R-Bench提供了一个系统化的诊断框架,用于评估视频世界模型是否能够弥合人至机器人之间的具身鸿沟,并将人类操作观察转化为以机器人为中心的训练资源。
English
Large-scale manipulation data is essential for robot learning, yet collecting robot demonstrations remains expensive and difficult to scale. Meanwhile, abundant egocentric human manipulation videos provide rich behavioral experiences, but transferring them across embodiments remains challenging due to differences between human hands and robotic end-effectors. Recent advances in video world models offer a promising pathway to synthesize robot-centric manipulation videos from human observations, while their cross-embodiment transfer capability remains largely unexplored. Therefore, we introduce H2R-Bench, a benchmark for evaluating cross-embodiment human-to-robot manipulation video generation, where models transform egocentric human demonstrations into robot manipulation videos under specified embodiments. Each benchmark instance contains a human demonstration video, target embodiment constraints, and source-grounded annotations covering task goals, action events, functional contacts, and object responses. H2R-Bench evaluates generated videos through five dimensions, including goal-state completion, action-event completion, functional contact transfer, embodiment correctness, and general video quality. We benchmark eleven state-of-the-art video generation models across six manipulation families and two robot embodiments. Our evaluation reveals that current video world models remain limited in human-to-robot manipulation transfer: even leading models often fail in embodiment consistency, functional interaction, and task execution. H2R-Bench provides a systematic diagnostic framework for evaluating whether video world models can bridge the human-to-robot embodiment gap and convert human manipulation observations into robot-centric training resources.