ForgeWM: 소수 스텝 행동 조건화 비디오 세계 모델을 위한 점진적 인과 훈련
ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models
August 14, 2026
저자: Xinye Li, Lingshuai Lin, Lei Wang, Liuzhou Zhang, Jialin Cui, Qingshan Li, Guanchu Wang, Qingbin Liu, Xi Chen, Jiang Bian, Wai Lam
cs.AI
초록
행동 조건부 비디오 세계 모델은 저지연 인과 생성과 게임 고유 컨트롤에 대한 신뢰할 수 있는 응답을 요구한다. 인과적 증류는 1단계 또는 소수 단계 비디오 합성을 가능하게 하지만, 이를 상호작용 세계 모델로 확장하는 것은 여전히 어려운데, 이산 키보드 상태와 연속 마우스 움직임이 인과 훈련 및 자기회귀 롤아웃 동안 시간적으로 압축된 잠재 청크와 정렬된 상태를 유지해야 하기 때문이다. 우리는 도메인 적응, 교사 강제 인과 훈련, 인과 일관성 증류, 그리고 양방향 교사와의 온폴리시 분포 정합을 통해 양방향 행동 조건부 비디오 생성기를 효율적인 소수 단계 세계 모델로 변환하는 점진적 프레임워크인 ForgeWM을 제안한다. 결과적으로 생성된 예산 특화 학생 모델은 1, 2, 4단계의 정상 상태 잡음 제거 예산으로 작동한다. ForgeWM은 또한 지연 시간에 민감한 상호작용과 선택적 재생 시간 정제를 결합한 이중 경로 배포 프로토콜을 지원하며, 여기서 1단계 학생 모델은 저장된 초안을 재잡음화하고 정제한다. 쌍으로 구성된 Minecraft 궤적에서 ForgeWM은 이미징 품질, 참조 정렬 모션 프로필 일치도, 행동 신호 정확도, 마우스 제어 정확도에서 평가된 시스템 중 최고를 기록했으며, 가장 낮은 참조 LPIPS를 달성했다. 동일한 4단계 절차는 게임패드로 제어되는 FPS 게임플레이에도 전이된다. 재생 시간 정제는 4단계 참조 품질에 도달하면서 잡음에서의 재생성보다 경험된 궤적에 약 3배 더 가깝게 유지된다. 이러한 결과는 제어 가능한 소수 단계 비디오 생성을 위한 ForgeWM의 효과성을 입증한다.
English
Action-conditioned video world models require low-latency causal generation and reliable responses to game-native controls. Although causal distillation enables one- or few-step video synthesis, extending it to interactive world models remains challenging, as discrete keyboard states and continuous mouse motion must remain aligned with temporally compressed latent chunks during causal training and autoregressive rollout. We introduce ForgeWM, a progressive framework that transforms a bidirectional action-conditioned video generator into efficient few-step world models through domain adaptation, teacher-forced causal training, causal consistency distillation, and on-policy distribution matching with a bidirectional teacher. The resulting budget-specialized students operate at steady-state denoising budgets of 1, 2, and 4 steps. ForgeWM further supports a dual-path deployment protocol combining latency-critical interaction with optional replay-time refinement, where the one-step student re-noises and refines its saved draft. On paired Minecraft trajectories, ForgeWM leads the evaluated systems in Imaging Quality, reference-aligned motion-profile agreement, action-sign accuracy, and mouse-control accuracy, while achieving the lowest reference LPIPS; the same four-stage recipe transfers to gamepad-controlled FPS gameplay. Replay-time refinement matches four-step reference quality while remaining roughly three times closer to the experienced trajectory than regeneration from noise. These results demonstrate ForgeWM's effectiveness for controllable few-step video generation.