ForgeWM:少步动作条件化视频世界模型的渐进式因果训练
ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models
August 14, 2026
作者: Xinye Li, Lingshuai Lin, Lei Wang, Liuzhou Zhang, Jialin Cui, Qingshan Li, Guanchu Wang, Qingbin Liu, Xi Chen, Jiang Bian, Wai Lam
cs.AI
摘要
以动作条件的视频世界模型需要低延迟的因果生成以及对游戏原生控制的可靠响应。尽管因果蒸馏能够实现单步或少步视频合成,但将其扩展到交互式世界模型仍具挑战性,因为在因果训练和自回归展开过程中,离散的键盘状态和连续的鼠标运动必须与时间压缩的潜在块保持对齐。我们提出了 ForgeWM,一个渐进式框架,通过域适应、教师强制的因果训练、因果一致性蒸馏以及基于策略的分布匹配(与双向教师)将双向动作条件视频生成器转化为高效的少步世界模型。由此得到的预算特化学生模型在稳态去噪预算下分别以 1、2 和 4 步运行。ForgeWM 还支持双路径部署协议,将延迟关键交互与可选的回放时细化相结合,其中单步学生模型对其保存的草稿进行重新加噪并细化。在配对的 Minecraft 轨迹上,ForgeWM 在成像质量、参考对齐的运动剖面一致性、动作信号准确性和鼠标控制准确性方面均领先于所评估的系统,同时实现了最低的参考 LPIPS;相同的四阶段方案可迁移至由游戏手柄控制的 FPS 游戏玩法。回放时细化能够匹配四步参考质量,同时与真实经历轨迹的距离约为从噪声重新生成的三分之一。这些结果证明了 ForgeWM 在可控少步视频生成方面的有效性。
English
Action-conditioned video world models require low-latency causal generation and reliable responses to game-native controls. Although causal distillation enables one- or few-step video synthesis, extending it to interactive world models remains challenging, as discrete keyboard states and continuous mouse motion must remain aligned with temporally compressed latent chunks during causal training and autoregressive rollout. We introduce ForgeWM, a progressive framework that transforms a bidirectional action-conditioned video generator into efficient few-step world models through domain adaptation, teacher-forced causal training, causal consistency distillation, and on-policy distribution matching with a bidirectional teacher. The resulting budget-specialized students operate at steady-state denoising budgets of 1, 2, and 4 steps. ForgeWM further supports a dual-path deployment protocol combining latency-critical interaction with optional replay-time refinement, where the one-step student re-noises and refines its saved draft. On paired Minecraft trajectories, ForgeWM leads the evaluated systems in Imaging Quality, reference-aligned motion-profile agreement, action-sign accuracy, and mouse-control accuracy, while achieving the lowest reference LPIPS; the same four-stage recipe transfers to gamepad-controlled FPS gameplay. Replay-time refinement matches four-step reference quality while remaining roughly three times closer to the experienced trajectory than regeneration from noise. These results demonstrate ForgeWM's effectiveness for controllable few-step video generation.