ChatPaper.aiChatPaper

ForgeWM:少步動作條件影片世界模型的漸進式因果訓練

ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models

August 14, 2026
作者: Xinye Li, Lingshuai Lin, Lei Wang, Liuzhou Zhang, Jialin Cui, Qingshan Li, Guanchu Wang, Qingbin Liu, Xi Chen, Jiang Bian, Wai Lam
cs.AI

摘要

動作條件的影片世界模型需要低延遲的因果生成,以及對遊戲原生控制項的可靠回應。儘管因果蒸餾能實現單步或少步的影片合成,但將其擴展至互動式世界模型仍具挑戰性,因為在因果訓練與自回歸展開期間,離散鍵盤狀態與連續滑鼠移動必須與時間壓縮的潛在區塊保持對齊。我們提出 ForgeWM,一個漸進式框架,透過域適應、教師強制因果訓練、因果一致性蒸餾,以及與雙向教師的同策略分佈匹配,將雙向動作條件的影片生成器轉化為高效率的少步世界模型。所產生的預算特化學生模型以 1、2、4 步的穩態去噪預算運作。ForgeWM 進一步支援雙路徑部署協定,結合延遲關鍵互動與可選的重放時細化,其中單步學生模型會對其儲存的草稿重新加噪並進行細化。在配對的 Minecraft 軌跡上,ForgeWM 在成像品質、參考對齊的運動輪廓一致性、動作訊號準確度與滑鼠控制準確度方面領先所有受評系統,同時達到最低的參考 LPIPS;相同的四階段配方也能遷移至使用遊戲手把控制的 FPS 遊戲。重放時細化可達到四步參考品質,同時與實際經歷軌跡的貼近程度約為從雜訊重新生成的三倍。這些結果證明了 ForgeWM 在可控少步影片生成上的有效性。
English
Action-conditioned video world models require low-latency causal generation and reliable responses to game-native controls. Although causal distillation enables one- or few-step video synthesis, extending it to interactive world models remains challenging, as discrete keyboard states and continuous mouse motion must remain aligned with temporally compressed latent chunks during causal training and autoregressive rollout. We introduce ForgeWM, a progressive framework that transforms a bidirectional action-conditioned video generator into efficient few-step world models through domain adaptation, teacher-forced causal training, causal consistency distillation, and on-policy distribution matching with a bidirectional teacher. The resulting budget-specialized students operate at steady-state denoising budgets of 1, 2, and 4 steps. ForgeWM further supports a dual-path deployment protocol combining latency-critical interaction with optional replay-time refinement, where the one-step student re-noises and refines its saved draft. On paired Minecraft trajectories, ForgeWM leads the evaluated systems in Imaging Quality, reference-aligned motion-profile agreement, action-sign accuracy, and mouse-control accuracy, while achieving the lowest reference LPIPS; the same four-stage recipe transfers to gamepad-controlled FPS gameplay. Replay-time refinement matches four-step reference quality while remaining roughly three times closer to the experienced trajectory than regeneration from noise. These results demonstrate ForgeWM's effectiveness for controllable few-step video generation.