マスクされた視覚行動による統一的世界モデリング
Masked Visual Actions for Unified World Modeling
July 21, 2026
著者: Hadi Alzayer, Wenlong Huang, Haonan Chen, Christopher Luey, Lvmin Zhang, Maneesh Agrawala, Gordon Wetzstein, Li Fei-Fei, Yilun Du, Jiajun Wu, Jia-Bin Huang
cs.AI
要旨
ビデオモデルは、視覚世界がどのように動き、相互作用し、接触に応答するかという豊富な事前知識を吸収しており、ロボットの世界モデリングのための有望な基盤となる。中心的な課題は、こうしたモデルに対して、それらが相互作用の事前知識を学習した視覚空間と整合しつつ、物理的な操作に根差した形で行動を伝達する方法にある。本稿では、マスク可視行動(Masked Visual Actions)を導入する。これはピクセル空間を介した制御インターフェースであり、行動をビデオ内の任意のエンティティの軌跡が部分的に明らかにされたものとして表現する。ロボットの動作を明らかにすることで、モデルは低レベルのロボット行動に対するシーンの応答を予測する前方動力学モデルとして機能する。一方、目的とする物体の動きを明らかにすることで、同じモデルがその結果と整合するロボットの振る舞いを回復する。実ビデオとシミュレーションからのマスク例をわずか15時間で微調整した単一のチェックポイントは、多様なシーンや複数の実施形態にわたって高い視覚的忠実性と制御可能性を達成する。下流の操作設定では、このモデルは、実際の実行結果と相関する想像上のロールアウトを生成し、政策評価を支援する。さらに、モデルベース計画における候補となる未来の順位付けにより意思決定を改善し、目的とする物体の動きからロボットの動作を合成する逆モデリングをサポートする。
English
Video models absorb rich priors over how the visual world moves, interacts, and responds to contact, making them promising substrates for robotic world modeling. The central challenge is how to communicate action to such models in a form aligned with the visual space in which they learned these interaction priors, yet still grounded in physical manipulation. We introduce Masked Visual Actions, a pixel-space control interface that expresses action as a partially revealed trajectory of an arbitrary entity in a video. Revealing robot motion makes the model act as a forward dynamics model that predicts the scene's response to low-level robot actions, while revealing desired object motion makes the same model recover robot behavior consistent with that outcome. Finetuned with only 15 hours of masked examples from real videos and simulation, a single checkpoint achieves strong visual fidelity and controllability across diverse scenes and multiple embodiments. In downstream manipulation settings, the model produces imagined rollouts whose outcomes correlate with real-world execution for policy evaluation, improves decision making by ranking candidate futures in model-based planning, and supports inverse modeling by synthesizing robot motion from desired object motion.