Zero-WAM:基于人类视频的上下文内世界-动作建模实现开放式任务泛化
Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization
August 26, 2026
作者: Jiaming Zhou, Qihang Zhang, Gangwei Xu, Cunxin Fan, Yujie Zhao, Ruilin Wang, Yiming Luo, Shuai Yang, Xing Zhu, Yujun Shen, Junwei Liang, Yinghao Xu
cs.AI
摘要
零样本跨任务泛化,即策略必须执行训练中从未见过的操作任务,仍然是机器人学习中的核心挑战。在大型语言模型中,只需在上下文中指定新任务即可执行,无需任何参数更新。这种上下文学习(ICL)形式将泛化转化为任务描述问题。为了实现跨任务泛化,我们将这一范式引入机器人操作,并提出操作任务的自然任务描述是人类视频:与语言不同,人类视频提供了关于预期任务演变的丰富视觉线索。我们提出了 Zero-WAM,一种因果视频-动作模型,通过遵循上下文中的人类视频引导来执行未见过的任务。为了解决任务丰富的配对人类-机器人数据稀缺的问题,我们提出了一种自动流程,将任务采样的机器人轨迹转换为语义匹配的人类视频,由此生成了 HumanGen 数据集,涵盖 8.6K 个任务下的 74.2K 个人类-机器人 ICL 对。在模型训练中,我们进一步引入了上下文未来分块预测(IFP)目标,该目标抑制从已见任务中学习到的捷径,并迫使策略从视频提示中提取任务信息。在 RoboTwin 2.0 仿真中的七个未见任务上,Zero-WAM 实现了 47.0% 的平均成功率,相比最强的视频-动作基线绝对提升了 29.5 个百分点。在真实世界评估中,它遵循人类视频引导,泛化到涉及多物体场景、长时域操作和精细插入的未见任务配置。
English
Zero-shot cross-task generalization, where a policy must execute manipulation tasks never seen during training, remains a central challenge in robot learning. In large language models, a novel task can be performed simply by specifying it in the context, without any parameter update. This form of in-context learning (ICL) turns generalization into a problem of task specification. To achieve cross-task generalization, we bring this paradigm to robotic manipulation, and argue that the natural task specification for manipulation is a human video: unlike language, it provides rich visual cues about the intended task evolution. We present Zero-WAM, a causal video-action model that executes unseen tasks by following in-context human video guidance. To address the scarcity of task-rich paired human-robot data, we propose an automatic pipeline that converts task-sampled robot trajectories into semantically matched human videos, yielding HumanGen, a dataset of 74.2K human-robot ICL pairs across 8.6K tasks. For model training, we further introduce an in-context future chunk prediction (IFP) objective that suppresses shortcuts learned from seen tasks and forces the policy to draw task information from the video prompt. On seven unseen tasks in RoboTwin 2.0 simulation, Zero-WAM achieves a 47.0% average success rate, an absolute improvement of 29.5 percentage points over the strongest video-action baseline. In real-world evaluations, it follows human video guidance to generalize to unseen task configurations involving multi-object scenes, long-horizon manipulation, and fine-grained insertion.