ChatPaper.aiChatPaper

ForgeWM: 少数ステップの行動条件付きビデオ世界モデルのための段階的因果学習

ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models

August 14, 2026
著者: Xinye Li, Lingshuai Lin, Lei Wang, Liuzhou Zhang, Jialin Cui, Qingshan Li, Guanchu Wang, Qingbin Liu, Xi Chen, Jiang Bian, Wai Lam
cs.AI

要旨

アクション条件付きビデオ世界モデルには、低レイテンシの因果的生成と、ゲームネイティブな操作に対する信頼性の高い応答が必要である。因果蒸留により1ステップまたは少数ステップの映像合成が可能になるが、これをインタラクティブな世界モデルに拡張することは依然として困難である。なぜなら、離散的なキーボード状態と連続的なマウス動作は、因果的訓練および自己回帰的ロールアウト中に、時間的に圧縮された潜在チャンクと整合し続けなければならないからである。我々は、ドメイン適応、教師強制による因果的訓練、因果整合性蒸留、および双方向教師とのオン方策分布マッチングを通じて、双方向アクション条件付き映像生成器を効率的な少数ステップ世界モデルへと変換する段階的フレームワークForgeWMを提案する。得られた予算特化型の生徒モデルは、1、2、4ステップの定常状態デノイジング予算で動作する。さらにForgeWMは、レイテンシが重要なインタラクションと任意のリプレイ時精緻化を組み合わせた二経路展開プロトコルをサポートしており、1ステップ生徒モデルが保存されたドラフトを再ノイズ化して精緻化する。ペア化されたMinecraft軌道において、ForgeWMは評価対象システム群の中で、画像品質、参照整合モーションプロファイル一致度、アクション符号精度、マウス制御精度で最高性能を示し、参照LPIPSは最低となった。同じ4段階の手法は、ゲームパッド操作のFPSゲームプレイにも転用できる。リプレイ時精緻化は4ステップ参照品質に匹敵し、ノイズからの再生成と比較して経験した軌道に対して約3倍近い。これらの結果は、制御可能な少数ステップ映像生成に対するForgeWMの有効性を示している。
English
Action-conditioned video world models require low-latency causal generation and reliable responses to game-native controls. Although causal distillation enables one- or few-step video synthesis, extending it to interactive world models remains challenging, as discrete keyboard states and continuous mouse motion must remain aligned with temporally compressed latent chunks during causal training and autoregressive rollout. We introduce ForgeWM, a progressive framework that transforms a bidirectional action-conditioned video generator into efficient few-step world models through domain adaptation, teacher-forced causal training, causal consistency distillation, and on-policy distribution matching with a bidirectional teacher. The resulting budget-specialized students operate at steady-state denoising budgets of 1, 2, and 4 steps. ForgeWM further supports a dual-path deployment protocol combining latency-critical interaction with optional replay-time refinement, where the one-step student re-noises and refines its saved draft. On paired Minecraft trajectories, ForgeWM leads the evaluated systems in Imaging Quality, reference-aligned motion-profile agreement, action-sign accuracy, and mouse-control accuracy, while achieving the lowest reference LPIPS; the same four-stage recipe transfers to gamepad-controlled FPS gameplay. Replay-time refinement matches four-step reference quality while remaining roughly three times closer to the experienced trajectory than regeneration from noise. These results demonstrate ForgeWM's effectiveness for controllable few-step video generation.