ChatPaper.aiChatPaper

StateFlow: 프리비주얼라이제이션을 위한 3D 세계 상태의 구축, 진화 및 접근

StateFlow: Building, Evolving, and Accessing 3D World States for Previsualization

August 12, 2026
저자: Yuyang Yin, Zixiang Li, Longxuan Deng, Hongkai Li, Shifang Zhao, Junnan Liu, Weirong Huang, Mengyu Wang, Tianxiao Fu, Yikai Wang, Peng-Shuai Wang, Xiaojie Jin, Yao Zhao, Yunchao Wei
cs.AI

초록

프리비주얼라이제이션은 영화, 게임, 건축 및 도시 설계에서 아이디어와 제작 사이의 중간 계층이다. 이를 통해 창작자는 장면, 동작, 카메라 및 시공간적 역학을 반복적으로 정교화할 수 있다. 그러나 기존 생성 방법은 원샷 이미지 또는 비디오 합성을 통해 이러한 모든 요소를 공동으로 제어하기 위해 단순한 프롬프트에 의존하므로, 제어 가능성이 약하고 반복적 편집에 대한 지원이 제한적이다. 근본적으로, 하나의 세계는 기하학, 외관 및 기타 속성을 가진 여러 요소와 카메라로 구성된다. 서로 다른 프레임은 대부분 재사용되는 공유 상태의 부분적 수정 또는 재조합을 통해 생성된다. 따라서 우리는 누락된 구성 요소가 명시적이고 지속적인 작업 상태라고 주장한다. 이를 해결하기 위해, 생성형 프리비주얼라이제이션을 위한 상태 중심 프레임워크인 StateFlow를 제시한다. StateFlow는 비디오를 원샷으로 생성하는 대신, 편집 가능한 3D 월드를 사용하여 장면 구조, 진화 및 카메라를 구성하며, 더 높은 충실도가 필요할 때 기성 비디오 모델이 시각적 품질을 향상시킨다. 이 월드는 장면 요소와 카메라 구성의 지속적인 구조화된 3D 상태로 유지되며, 프리비주얼라이제이션의 핵심 작업 표현 역할을 한다. 이러한 통찰력을 바탕으로 StateFlow는 월드 상태를 구축, 진화 및 접근하는 세 가지 단계로 구성된다. 상태 구축은 사전 기반의 충돌 인식 이중 뷰 초기화를 통해 생성된 2D 콘텐츠를 일관된 3D 월드로 승격시키며, 상태 진화는 월드 메모리를 보존하면서 사용자 의도를 구조화된 상태 전이로 변환하여 각 편집에 대해 전체 장면 재생성을 피한다. 상태 접근은 렌더 피드백 기반 성찰을 통해 카메라 계획을 시각적으로 실행 가능한 궤적으로 정제함으로써 VLM 의미론에만 의존하지 않는다. 실험 결과는 StateFlow가 비디오 제작 및 게임형 프로토타이핑을 위한 고품질 3D 월드를 생성함을 보여준다.
English
Previsualization is an intermediate layer between ideas and production in film, games, architecture, and urban design. It lets creators iteratively refine scenes, actions, cameras, and spatial-temporal dynamics. Yet existing generative methods rely on simple prompts to jointly control all of these factors through one-shot image or video synthesis, offering weak controllability and limited support for iterative editing. Fundamentally, a world comprises multiple elements with geometry, appearance, and other attributes, together with cameras. Different frames are produced through local modifications or recombinations of this shared state, which is otherwise largely reused. Therefore, we argue that the missing component is an explicit and persistent working state. To address this, we present StateFlow, a state-centric framework for generative previsualization. Rather than generating videos in one shot, StateFlow uses an editable 3D world to organize scene structure, evolution, and cameras, while off-the-shelf video models enhance visual quality when higher fidelity is desired. This world is maintained as a persistent structured 3D state of scene elements and camera configurations, serving as the core working representation for previsualization. Built on this insight, StateFlow has three stages to construct, evolve, and access the world state. State construction lifts generated 2D content into a coherent 3D world through prior-guided, conflict-aware dual-view initialization, while State evolution translates user intent into structured state transitions while preserving world memory, avoiding full-scene regeneration for each edit. State access uses render-feedback reflection to refine camera plans into visually feasible trajectories, avoiding reliance on VLM semantics alone. Experiments show that StateFlow produces high-quality 3D worlds for video creation and game-like prototyping.