Zero-WAM:從人類影片中進行情境內世界-動作建模以實現開放式任務泛化
Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization
August 26, 2026
作者: Jiaming Zhou, Qihang Zhang, Gangwei Xu, Cunxin Fan, Yujie Zhao, Ruilin Wang, Yiming Luo, Shuai Yang, Xing Zhu, Yujun Shen, Junwei Liang, Yinghao Xu
cs.AI
摘要
零樣本跨任務泛化——即策略必須執行訓練期間從未見過的操作任務——仍是機器人學習中的核心挑戰。在大語言模型中,一項新任務只需在上下文中指定即可執行,無需任何參數更新。這種形式的上下文學習(in-context learning, ICL)將泛化轉化為任務規格的問題。為了實現跨任務泛化,我們將此範式引入機器人操作,並主張操作任務的自然規格是人類影片:與語言不同,人類影片提供了關於預期任務演變的豐富視覺線索。我們提出 Zero-WAM,這是一個因果影片-動作模型,透過遵循以人類影片為上下文的引導來執行未見過的任務。為了解決任務豐富的人類-機器人配對資料稀缺的問題,我們提出一個自動化流程,將任務取樣所得的機器人軌跡轉換為語義匹配的人類影片,從而生成 HumanGen 資料集,其中包含 8.6K 個任務的 74.2K 個人類-機器人 ICL 配對。在模型訓練方面,我們進一步引入了上下文未來片段預測(in-context future chunk prediction, IFP)目標,以抑制從已見任務中學到的捷徑,並迫使策略從影片提示中提取任務資訊。在 RoboTwin 2.0 模擬環境的七個未見過任務上,Zero-WAM 達到了 47.0% 的平均成功率,比最強的影片-動作基線方法絕對提升了 29.5 個百分點。在真實世界評估中,它能遵循人類影片引導,泛化到涉及多物體場景、長程操作和精細插入的未見過任務配置。
English
Zero-shot cross-task generalization, where a policy must execute manipulation tasks never seen during training, remains a central challenge in robot learning. In large language models, a novel task can be performed simply by specifying it in the context, without any parameter update. This form of in-context learning (ICL) turns generalization into a problem of task specification. To achieve cross-task generalization, we bring this paradigm to robotic manipulation, and argue that the natural task specification for manipulation is a human video: unlike language, it provides rich visual cues about the intended task evolution. We present Zero-WAM, a causal video-action model that executes unseen tasks by following in-context human video guidance. To address the scarcity of task-rich paired human-robot data, we propose an automatic pipeline that converts task-sampled robot trajectories into semantically matched human videos, yielding HumanGen, a dataset of 74.2K human-robot ICL pairs across 8.6K tasks. For model training, we further introduce an in-context future chunk prediction (IFP) objective that suppresses shortcuts learned from seen tasks and forces the policy to draw task information from the video prompt. On seven unseen tasks in RoboTwin 2.0 simulation, Zero-WAM achieves a 47.0% average success rate, an absolute improvement of 29.5 percentage points over the strongest video-action baseline. In real-world evaluations, it follows human video guidance to generalize to unseen task configurations involving multi-object scenes, long-horizon manipulation, and fine-grained insertion.