用於統一世界建模的遮罩視覺動作
Masked Visual Actions for Unified World Modeling
July 21, 2026
作者: Hadi Alzayer, Wenlong Huang, Haonan Chen, Christopher Luey, Lvmin Zhang, Maneesh Agrawala, Gordon Wetzstein, Li Fei-Fei, Yilun Du, Jiajun Wu, Jia-Bin Huang
cs.AI
摘要
视频模型吸收了关于视觉世界如何运动、交互以及响应接触的丰富先验知识,使其成为机器人世界建模的理想基底。核心挑战在于如何将动作传达给这类模型——既需要以与其学习交互先验的视觉空间一致的形式表达,又必须植根于物理操作。我们提出掩码视觉动作(Masked Visual Actions),一种像素空间控制接口,将动作表示为视频中任意实体部分显现的运动轨迹。揭示机器人运动可使模型充当正向动力学模型,预测场景对低级机器人动作的响应;而揭示期望的物体运动则能使同一模型恢复与该结果一致的机器人行为。仅通过15小时来自真实视频和模拟的掩码示例进行微调,单个检查点便能在多样化场景和多形态实体中实现高视觉保真度与可控性。在下游操作场景中,该模型可生成与真实世界执行结果相关联的想象展开序列,用于策略评估;通过基于模型的规划中对候选未来状态进行排序来改进决策;并支持逆向建模,从期望的物体运动合成机器人运动。
English
Video models absorb rich priors over how the visual world moves, interacts, and responds to contact, making them promising substrates for robotic world modeling. The central challenge is how to communicate action to such models in a form aligned with the visual space in which they learned these interaction priors, yet still grounded in physical manipulation. We introduce Masked Visual Actions, a pixel-space control interface that expresses action as a partially revealed trajectory of an arbitrary entity in a video. Revealing robot motion makes the model act as a forward dynamics model that predicts the scene's response to low-level robot actions, while revealing desired object motion makes the same model recover robot behavior consistent with that outcome. Finetuned with only 15 hours of masked examples from real videos and simulation, a single checkpoint achieves strong visual fidelity and controllability across diverse scenes and multiple embodiments. In downstream manipulation settings, the model produces imagined rollouts whose outcomes correlate with real-world execution for policy evaluation, improves decision making by ranking candidate futures in model-based planning, and supports inverse modeling by synthesizing robot motion from desired object motion.