Zero-WAM: 人間のビデオからのインコンテキスト世界行動モデリングによるオープンエンドなタスク汎化
Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization
August 26, 2026
著者: Jiaming Zhou, Qihang Zhang, Gangwei Xu, Cunxin Fan, Yujie Zhao, Ruilin Wang, Yiming Luo, Shuai Yang, Xing Zhu, Yujun Shen, Junwei Liang, Yinghao Xu
cs.AI
要旨
ゼロショットのクロスタスク汎化、すなわちポリシーが訓練中に一度も見たことのない操作タスクを実行しなければならない問題は、ロボット学習における中心的な課題であり続けている。大規模言語モデルでは、新しいタスクはパラメータ更新を一切行わずに、コンテキスト内で指定するだけで実行できる。このインコンテキスト学習(ICL)の形式は、汎化をタスク指定の問題へと転換する。クロスタスク汎化を達成するために、我々はこのパラダイムをロボット操作に導入し、操作における自然なタスク指定は人間のビデオであると主張する。言語とは異なり、ビデオは意図されたタスクの進展に関する豊かな視覚的手がかりを提供するからである。我々は、インコンテキストの人間ビデオガイダンスに従って未見タスクを実行する因果ビデオアクションモデル、Zero-WAMを提案する。タスクに富んだペアの人間-ロボットデータの不足に対処するため、我々はタスクサンプリングされたロボット軌道を意味的に整合した人間ビデオへ変換する自動パイプラインを提案し、8.6Kタスクにわたる74.2Kの人間-ロボットICLペアからなるデータセットHumanGenを生成する。モデル訓練のために、さらに、既知タスクから学習されるショートカットを抑制し、ポリシーがビデオプロンプトからタスク情報を引き出すことを強制する、インコンテキスト未来チャンク予測(IFP)目的関数を導入する。RoboTwin 2.0シミュレーションにおける7つの未見タスクでは、Zero-WAMは平均成功率47.0%を達成し、最強のビデオアクションベースラインに対して29.5パーセントポイントの絶対的な改善を示す。実世界評価では、人間のビデオガイダンスに従い、複数物体シーン、長期にわたる操作、高精度な挿入を含む未見のタスク構成へ汎化する。
English
Zero-shot cross-task generalization, where a policy must execute manipulation tasks never seen during training, remains a central challenge in robot learning. In large language models, a novel task can be performed simply by specifying it in the context, without any parameter update. This form of in-context learning (ICL) turns generalization into a problem of task specification. To achieve cross-task generalization, we bring this paradigm to robotic manipulation, and argue that the natural task specification for manipulation is a human video: unlike language, it provides rich visual cues about the intended task evolution. We present Zero-WAM, a causal video-action model that executes unseen tasks by following in-context human video guidance. To address the scarcity of task-rich paired human-robot data, we propose an automatic pipeline that converts task-sampled robot trajectories into semantically matched human videos, yielding HumanGen, a dataset of 74.2K human-robot ICL pairs across 8.6K tasks. For model training, we further introduce an in-context future chunk prediction (IFP) objective that suppresses shortcuts learned from seen tasks and forces the policy to draw task information from the video prompt. On seven unseen tasks in RoboTwin 2.0 simulation, Zero-WAM achieves a 47.0% average success rate, an absolute improvement of 29.5 percentage points over the strongest video-action baseline. In real-world evaluations, it follows human video guidance to generalize to unseen task configurations involving multi-object scenes, long-horizon manipulation, and fine-grained insertion.