ChatPaper.aiChatPaper

Stream4D: 스트리밍 자기회귀 확산 비디오 모델을 위한 4D 일관성

Stream4D: 4D-Consistency for Streaming Autoregressive Diffusion Video Models

August 20, 2026
저자: Yuanhao Ban, Jiaqi Feng, Hengguang Zhou, Xiaohuan Pei, Justin Cui, Cho-Jui Hsieh
cs.AI

초록

스트리밍 자기회귀 확산 모델은 실시간 장기 비디오 생성을 가능하게 하지만, 그 학습 목표는 일관된 세계의 기하학과 역학보다는 로컬 프레임 예측을 최적화한다. 그 결과 긴 롤아웃은 기하학적 드리프트를 누적시키고 정적이거나 부자연스러운 움직임으로 퇴화된다. 최근 양방향 접근법들은 3D 가우시안 스플래팅 재구성에 기반한 보상 신호를 사용하여 이 문제를 해결하고자 한다. 그러나 단일 강체 3D 재구성은 동적 장면을 모델링할 수 없으므로, 이 비평자는 실제 객체 움직임을 재구성 오류로 페널티를 부과하며, 비디오를 정지시킴으로써 최대화된다. 이러한 지름길은 각 청크가 이미 정적인 구성을 전파할 수 있는 AR(자기회귀) 환경에서 특히 해롭다. 본 연구에서는 정적 비평자를 장면 역학을 명시적으로 모델링하는 피드포워드 4D 재구성 보상으로 대체하여, 일관된 움직임이 높은 일관성 보상을 받을 수 있게 하는 Stream4D를 제안한다. 또한 움직임의 크기와 품질을 추가로 안내하기 위해, 자연스러운 장면 흐름 크기에 보상을 부여하고 지터 및 비강체 아티팩트에는 페널티를 부과하는 움직임 사전을 추가한다. 최종 레시피는 이 두 항목을 경량 지각 앵커와 결합한다. 다양한 자기회귀 비디오 백본과 다양한 생성 지평에 걸쳐, Stream4D는 4D 재구성 품질을 개선하고, 움직임을 더 효과적으로 보존하며, 더 높은 인간 정렬 선호도를 달성한다. 프로젝트 페이지: https://banyuanhao.github.io/Stream4D/
English
Streaming autoregressive diffusion models enable real-time, long-horizon video generation, but their training objectives optimize local frame prediction rather than the geometry and dynamics of a coherent world: long rollouts accumulate geometric drift and degrade into static or unnatural motion. Recent bidirectional approaches address this problem using rewards signals built upon 3D Gaussian-Splatting reconstruction. However, a single rigid 3d reconstruction cannot model a dynamic scene, so this critic penalizes genuine object motion as reconstruction error and is maximized by freezing the video. This shortcut is especially detrimental in the AR setting, where each chunk can propagate an already-static configuration. In this work, we propose Stream4D, which replaces the static critic with a feed-forward 4D reconstruction reward that explicitly models scene dynamics, allowing coherent motion to receive high consistency rewards. To further guide motion magnitude and quality, we add a motion prior that rewards natural scene-flow magnitude while penalizing jitter and non-rigid artifacts. Our final recipe combines these two terms with a lightweight perceptual anchor. Across various autoregressive video backbones and various generation horizons, Stream4D improves 4D reconstruction quality, preserves motion more effectively, and achieves higher human-aligned preference. Project page: https://banyuanhao.github.io/Stream4D/