ChatPaper.aiChatPaper

Flex-Forcing:邁向統一的自迴歸與雙向視頻擴散模型

Flex-Forcing: Towards a Unified Autoregressive and Bidirectional Video Diffusion Model

July 3, 2026
作者: Xinyin Ma, Julius Berner, Chao Liu, Arash Vahdat, Weili Nie, Xinchao Wang
cs.AI

摘要

大規模生成式模型的最新進展已大幅推進了影片生成技術,然而現有方法仍受限於僵硬的推理範式。雙向擴散模型在全局連貫性與視覺保真度上表現優異,但推理速度緩慢;而自回歸模型則能提供高效且串流式的生成,卻以長程一致性與曝光偏差為代價。我們提出Flex-Forcing,一個統一的訓練與推理框架,使影片擴散模型能同時在雙向與自回歸生成模式下無縫運作。其核心概念是一種沿時間軸與去噪步驟共同定義的靈活分塊機制。此設計使模型能夠:(1) 根據不同裝置的運算預算執行靈活分塊;(2) 跨分塊進行雙向推理以規劃全局結構,同時在每個分塊內部以自回歸方式生成影格,達成高效且細粒度的合成;(3) 在無嚴格因果限制下,實現任意順序、任意時間步的自回歸生成。在多項影片生成基準上的大量實驗顯示,相較於採用固定推理排程的強基線模型,Flex-Forcing能持續產出更佳的影片品質與長片穩定性,同時提供更快的推理速度。
English
Recent progress in large-scale generative models has substantially advanced video generation, yet existing methods remain constrained by a rigid inference paradigm. Bidirectional diffusion models excel at global coherence and visual fidelity but suffer from slow inference, while autoregressive models offer efficient and streaming generation at the cost of long-range consistency and exposure bias. We introduce Flex-Forcing, a unified training and inference framework that enables a video diffusion model to seamlessly operate under both bidirectional and autoregressive generation regimes. The core idea is a flexible chunking mechanism jointly defined over the temporal axis and denoising steps. This design allows the model to (1) perform flexible chunking according to different device budgets, (2) perform bidirectional inference across chunks for global structure planning, while generating frames autoregressively within each chunk for efficient and fine-grained synthesis, and (3) perform any-order, any-timestep autoregressive generation without the strict causal constraint. Extensive experiments on multiple video generation benchmarks demonstrate that Flex-Forcing achieves consistently better video quality, long-video stability than strong baselines with a rigid inference schedule, while offering faster inference.