ChatPaper.aiChatPaper

통합 세계 모델링을 위한 마스킹된 시각적 행동

Masked Visual Actions for Unified World Modeling

July 21, 2026
저자: Hadi Alzayer, Wenlong Huang, Haonan Chen, Christopher Luey, Lvmin Zhang, Maneesh Agrawala, Gordon Wetzstein, Li Fei-Fei, Yilun Du, Jiajun Wu, Jia-Bin Huang
cs.AI

초록

비디오 모델은 시각적 세계가 어떻게 움직이고, 상호작용하며, 접촉에 반응하는지에 대한 풍부한 사전 지식을 흡수하여, 로봇 세계 모델링을 위한 유망한 기반이 된다. 핵심 과제는 이러한 상호작용 사전 지식을 학습한 시각적 공간과 정렬되면서도 물리적 조작에 기반한 형태로, 그러한 모델에 행동을 전달하는 방법이다. 우리는 마스크된 시각적 행동(Masked Visual Actions)을 제안한다. 이는 비디오 내 임의 개체의 부분적으로 드러난 궤적으로 행동을 표현하는 픽셀 공간 제어 인터페이스이다. 로봇 움직임을 드러내면 모델은 저수준 로봇 행동에 대한 장면의 반응을 예측하는 순방향 동역학 모델로 작동하고, 원하는 객체 움직임을 드러내면 동일한 모델이 그 결과와 일치하는 로봇 행동을 복원한다. 실제 비디오와 시뮬레이션에서 단 15시간의 마스크된 예제로 미세 조정된 단일 체크포인트는 다양한 장면과 여러 형체에서 강력한 시각적 충실도와 제어 가능성을 달성한다. 하위 조작 설정에서, 이 모델은 정책 평가를 위해 실제 실행과 결과가 상관관계를 가지는 가상 롤아웃을 생성하고, 모델 기반 계획에서 후보 미래를 순위화하여 의사 결정을 개선하며, 원하는 객체 움직임으로부터 로봇 움직임을 합성하는 역모델링을 지원한다.
English
Video models absorb rich priors over how the visual world moves, interacts, and responds to contact, making them promising substrates for robotic world modeling. The central challenge is how to communicate action to such models in a form aligned with the visual space in which they learned these interaction priors, yet still grounded in physical manipulation. We introduce Masked Visual Actions, a pixel-space control interface that expresses action as a partially revealed trajectory of an arbitrary entity in a video. Revealing robot motion makes the model act as a forward dynamics model that predicts the scene's response to low-level robot actions, while revealing desired object motion makes the same model recover robot behavior consistent with that outcome. Finetuned with only 15 hours of masked examples from real videos and simulation, a single checkpoint achieves strong visual fidelity and controllability across diverse scenes and multiple embodiments. In downstream manipulation settings, the model produces imagined rollouts whose outcomes correlate with real-world execution for policy evaluation, improves decision making by ranking candidate futures in model-based planning, and supports inverse modeling by synthesizing robot motion from desired object motion.