Zero-WAM: 개방형 작업 일반화를 위한 인간 영상 기반 인컨텍스트 세계-행동 모델링
Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization
August 26, 2026
저자: Jiaming Zhou, Qihang Zhang, Gangwei Xu, Cunxin Fan, Yujie Zhao, Ruilin Wang, Yiming Luo, Shuai Yang, Xing Zhu, Yujun Shen, Junwei Liang, Yinghao Xu
cs.AI
초록
제로샷 교차 작업 일반화, 즉 훈련 중에 본 적 없는 조작 작업을 정책이 실행해야 하는 문제는 로봇 학습의 핵심 과제로 남아 있다. 대규모 언어 모델에서는 매개변수 갱신 없이 맥락에서 새로운 작업을 지정하기만 해도 수행할 수 있다. 이러한 형태의 맥락 내 학습(ICL)은 일반화를 작업 지정의 문제로 전환한다. 교차 작업 일반화를 달성하기 위해 우리는 이 패러다임을 로봇 조작에 적용하며, 조작 작업의 자연스러운 지정 방식은 인간 비디오라고 주장한다. 언어와 달리 인간 비디오는 의도된 작업 진행에 대한 풍부한 시각적 단서를 제공한다. 우리는 맥락 내 인간 비디오 안내를 따라 미지의 작업을 실행하는 인과적 비디오-행동 모델인 Zero-WAM을 제시한다. 작업이 풍부한 인간-로봇 짝 데이터의 부족을 해결하기 위해, 작업 표본 로봇 궤적을 의미적으로 대응하는 인간 비디오로 변환하는 자동 파이프라인을 제안하며, 이를 통해 8.6K 작업에 걸쳐 74.2K 개의 인간-로봇 ICL 쌍으로 구성된 데이터셋 HumanGen을 구축한다. 모델 훈련을 위해, 우리는 추가로 기존 작업에서 학습된 지름길을 억제하고 정책이 비디오 프롬프트에서 작업 정보를 얻도록 강제하는 맥락 내 미래 청크 예측(IFP) 목적 함수를 도입한다. RoboTwin 2.0 시뮬레이션의 7가지 미지의 작업에서 Zero-WAM은 평균 성공률 47.0%를 달성하여, 가장 강력한 비디오-행동 기준 모델 대비 절대 29.5%포인트 향상되었다. 실제 환경 평가에서, Zero-WAM은 인간 비디오 안내를 따라 다중 객체 장면, 장기 조작, 정밀 삽입을 포함한 미지의 작업 구성으로 일반화한다.
English
Zero-shot cross-task generalization, where a policy must execute manipulation tasks never seen during training, remains a central challenge in robot learning. In large language models, a novel task can be performed simply by specifying it in the context, without any parameter update. This form of in-context learning (ICL) turns generalization into a problem of task specification. To achieve cross-task generalization, we bring this paradigm to robotic manipulation, and argue that the natural task specification for manipulation is a human video: unlike language, it provides rich visual cues about the intended task evolution. We present Zero-WAM, a causal video-action model that executes unseen tasks by following in-context human video guidance. To address the scarcity of task-rich paired human-robot data, we propose an automatic pipeline that converts task-sampled robot trajectories into semantically matched human videos, yielding HumanGen, a dataset of 74.2K human-robot ICL pairs across 8.6K tasks. For model training, we further introduce an in-context future chunk prediction (IFP) objective that suppresses shortcuts learned from seen tasks and forces the policy to draw task information from the video prompt. On seven unseen tasks in RoboTwin 2.0 simulation, Zero-WAM achieves a 47.0% average success rate, an absolute improvement of 29.5 percentage points over the strongest video-action baseline. In real-world evaluations, it follows human video guidance to generalize to unseen task configurations involving multi-object scenes, long-horizon manipulation, and fine-grained insertion.