ChatPaper.aiChatPaper

Flex-Forcing: 統一的自己回帰・双方向ビデオ拡散モデルに向けて

Flex-Forcing: Towards a Unified Autoregressive and Bidirectional Video Diffusion Model

July 3, 2026
著者: Xinyin Ma, Julius Berner, Chao Liu, Arash Vahdat, Weili Nie, Xinchao Wang
cs.AI

要旨

大規模生成モデルの最近の進歩により、動画生成は大幅に向上したが、既存の手法は依然として固定された推論パラダイムに制約されている。双方向拡散モデルは全体的な一貫性と視覚的忠実度に優れるものの、推論速度が遅く、一方、自己回帰モデルは効率的でストリーミング的な生成を実現するが、長距離の一貫性や露出バイアスに課題がある。そこで我々は、ビデオ拡散モデルが双方向生成と自己回帰生成の両方の枠組みでシームレスに動作することを可能にする、統一的な学習・推論フレームワークであるFlex-Forcingを提案する。その核となるアイデアは、時間軸とノイズ除去ステップの両方にわたって共同で定義される柔軟なチャンク分割機構である。この設計により、モデルは (1) 異なるデバイス予算に応じて柔軟にチャンク分割を実行し、(2) 全体的な構造計画のためにチャンク間で双方向推論を行いつつ、各チャンク内ではフレームを自己回帰的に生成して効率的かつ詳細な合成を実現し、(3) 厳密な因果的制約なしに任意の順序・任意のタイムステップでの自己回帰生成を実行できる。複数の動画生成ベンチマークにおける広範な実験により、Flex-Forcingは固定された推論スケジュールを持つ強力なベースラインと比較して、一貫して優れた動画品質と長尺動画の安定性を達成し、かつ高速な推論を提供することを実証する。
English
Recent progress in large-scale generative models has substantially advanced video generation, yet existing methods remain constrained by a rigid inference paradigm. Bidirectional diffusion models excel at global coherence and visual fidelity but suffer from slow inference, while autoregressive models offer efficient and streaming generation at the cost of long-range consistency and exposure bias. We introduce Flex-Forcing, a unified training and inference framework that enables a video diffusion model to seamlessly operate under both bidirectional and autoregressive generation regimes. The core idea is a flexible chunking mechanism jointly defined over the temporal axis and denoising steps. This design allows the model to (1) perform flexible chunking according to different device budgets, (2) perform bidirectional inference across chunks for global structure planning, while generating frames autoregressively within each chunk for efficient and fine-grained synthesis, and (3) perform any-order, any-timestep autoregressive generation without the strict causal constraint. Extensive experiments on multiple video generation benchmarks demonstrate that Flex-Forcing achieves consistently better video quality, long-video stability than strong baselines with a rigid inference schedule, while offering faster inference.