Flex-Forcing: 통합된 자기회귀 및 양방향 비디오 확산 모델을 향하여
Flex-Forcing: Towards a Unified Autoregressive and Bidirectional Video Diffusion Model
July 3, 2026
저자: Xinyin Ma, Julius Berner, Chao Liu, Arash Vahdat, Weili Nie, Xinchao Wang
cs.AI
초록
대규모 생성 모델의 최근 진전은 비디오 생성을 상당히 발전시켰지만, 기존 방법들은 여전히 경직된 추론 패러다임에 의해 제약을 받는다. 양방향 확산 모델은 전역적 일관성과 시각적 충실도에 뛰어나지만 추론 속도가 느린 반면, 자기회귀 모델은 효율적이고 스트리밍 방식의 생성을 제공하지만 장거리 일관성과 노출 편향을 희생한다. 우리는 Flex-Forcing을 소개한다. 이는 비디오 확산 모델이 양방향 및 자기회귀 생성 방식 모두에서 원활하게 작동할 수 있도록 하는 통합 훈련 및 추론 프레임워크이다. 핵심 아이디어는 시간 축과 노이즈 제거 단계에 걸쳐 공동으로 정의된 유연한 청킹 메커니즘이다. 이 설계는 모델이 (1) 다양한 장치 예산에 따라 유연한 청킹을 수행하고, (2) 전역 구조 계획을 위해 청크 간 양방향 추론을 수행하면서 각 청크 내에서는 프레임을 자기회귀적으로 생성하여 효율적이고 세밀한 합성을 수행하며, (3) 엄격한 인과적 제약 없이 임의 순서 및 임의 시간 단계의 자기회귀 생성을 수행할 수 있도록 한다. 여러 비디오 생성 벤치마크에 대한 광범위한 실험 결과, Flex-Forcing이 경직된 추론 일정을 가진 강력한 기준선보다 일관되게 더 나은 비디오 품질과 긴 비디오 안정성을 달성하면서 더 빠른 추론을 제공함을 보여준다.
English
Recent progress in large-scale generative models has substantially advanced video generation, yet existing methods remain constrained by a rigid inference paradigm. Bidirectional diffusion models excel at global coherence and visual fidelity but suffer from slow inference, while autoregressive models offer efficient and streaming generation at the cost of long-range consistency and exposure bias. We introduce Flex-Forcing, a unified training and inference framework that enables a video diffusion model to seamlessly operate under both bidirectional and autoregressive generation regimes. The core idea is a flexible chunking mechanism jointly defined over the temporal axis and denoising steps. This design allows the model to (1) perform flexible chunking according to different device budgets, (2) perform bidirectional inference across chunks for global structure planning, while generating frames autoregressively within each chunk for efficient and fine-grained synthesis, and (3) perform any-order, any-timestep autoregressive generation without the strict causal constraint. Extensive experiments on multiple video generation benchmarks demonstrate that Flex-Forcing achieves consistently better video quality, long-video stability than strong baselines with a rigid inference schedule, while offering faster inference.