ChatPaper.aiChatPaper

Stream4D:串流自迴歸擴散影片模型的4D一致性

Stream4D: 4D-Consistency for Streaming Autoregressive Diffusion Video Models

August 20, 2026
作者: Yuanhao Ban, Jiaqi Feng, Hengguang Zhou, Xiaohuan Pei, Justin Cui, Cho-Jui Hsieh
cs.AI

摘要

串流自迴歸擴散模型實現了即時、長時程的視訊生成,但其訓練目標僅最佳化局部幀預測,而非一個連貫世界的幾何與動態:長時間的展開會累積幾何漂移,並退化為靜態或非自然的運動。近期的雙向方法利用基於3D高斯潑濺重建的獎勵訊號來解決此問題。然而,單一剛性3D重建無法建模動態場景,因此該評論器會將真實的物體運動視為重建誤差而加以懲罰,且可透過凍結視訊將其最大化。此捷徑在自迴歸(AR)設定中尤其有害,因為每個區塊都可能傳播已靜止的配置。在本工作中,我們提出Stream4D,以明確建模場景動態的前饋4D重建獎勵取代靜態評論器,使連貫運動能獲得高一致性獎勵。為進一步引導運動幅度與品質,我們加入一項運動先驗,獎勵自然的場景流幅度,同時懲罰抖動與非剛性偽影。我們最終的方案結合了這兩項以及一個輕量級感知錨點。在各種自迴歸視訊骨幹網路與各種生成時程下,Stream4D提升了4D重建品質、更有效地保留運動,並實現更高的人類對齊偏好度。專案頁面:https://banyuanhao.github.io/Stream4D/
English
Streaming autoregressive diffusion models enable real-time, long-horizon video generation, but their training objectives optimize local frame prediction rather than the geometry and dynamics of a coherent world: long rollouts accumulate geometric drift and degrade into static or unnatural motion. Recent bidirectional approaches address this problem using rewards signals built upon 3D Gaussian-Splatting reconstruction. However, a single rigid 3d reconstruction cannot model a dynamic scene, so this critic penalizes genuine object motion as reconstruction error and is maximized by freezing the video. This shortcut is especially detrimental in the AR setting, where each chunk can propagate an already-static configuration. In this work, we propose Stream4D, which replaces the static critic with a feed-forward 4D reconstruction reward that explicitly models scene dynamics, allowing coherent motion to receive high consistency rewards. To further guide motion magnitude and quality, we add a motion prior that rewards natural scene-flow magnitude while penalizing jitter and non-rigid artifacts. Our final recipe combines these two terms with a lightweight perceptual anchor. Across various autoregressive video backbones and various generation horizons, Stream4D improves 4D reconstruction quality, preserves motion more effectively, and achieves higher human-aligned preference. Project page: https://banyuanhao.github.io/Stream4D/