StateFlow: プリビジュアライゼーションのための3Dワールド状態の構築、進化、およびアクセス
StateFlow: Building, Evolving, and Accessing 3D World States for Previsualization
August 12, 2026
著者: Yuyang Yin, Zixiang Li, Longxuan Deng, Hongkai Li, Shifang Zhao, Junnan Liu, Weirong Huang, Mengyu Wang, Tianxiao Fu, Yikai Wang, Peng-Shuai Wang, Xiaojie Jin, Yao Zhao, Yunchao Wei
cs.AI
要旨
プリビジュアライゼーションは、映画、ゲーム、建築、都市デザインにおけるアイデアと制作の間の中間層である。これによりクリエイターは、シーン、アクション、カメラ、時空間ダイナミクスを反復的に洗練できる。しかし既存の生成手法は、単純なプロンプトに依存してこれらすべての要素を一度の画像・映像合成で統合制御するものであり、制御性が弱く、反復的編集のサポートも限定的である。本質的に、ワールドはジオメトリ、外観、その他の属性を持つ複数の要素とカメラで構成される。異なるフレームは、この共有状態の局所的な変更または再結合によって生成されるが、それ以外の部分はほぼ再利用される。したがって我々は、欠けている要素は明示的かつ永続的な作業状態(ワーキングステート)であると主張する。この課題に対処するため、本稿では生成的プリビジュアライゼーションのための状態中心フレームワークであるStateFlowを提案する。StateFlowは、映像を一度に生成するのではなく、編集可能な3Dワールドを用いてシーン構造・進化・カメラを整理し、より高い忠実度が望まれる場合には既存の映像モデルが画質を向上させる。このワールドは、シーン要素とカメラ構成からなる永続的な構造化3D状態として維持され、プリビジュアライゼーションの中核となる作業表現として機能する。この洞察に基づき、StateFlowはワールド状態の構築・進化・アクセスという3つの段階を持つ。状態構築は、事前知識に導かれ競合を考慮したデュアルビュー初期化により、生成された2Dコンテンツを一貫性のある3Dワールドへとリフトアップする。状態進化は、ワールドメモリを保持しつつユーザーの意図を構造化された状態遷移へ変換し、編集のたびにシーン全体を再生成することを回避する。状態アクセスは、レンダーフィードバックに基づく内省的プロセスによってカメラプランを視覚的に実現可能な軌道へと洗練し、VLMの意味論のみに依存することを避ける。実験により、StateFlowは映像制作やゲーム的プロトタイピングのための高品質な3Dワールドを生成することが示された。
English
Previsualization is an intermediate layer between ideas and production in film, games, architecture, and urban design. It lets creators iteratively refine scenes, actions, cameras, and spatial-temporal dynamics. Yet existing generative methods rely on simple prompts to jointly control all of these factors through one-shot image or video synthesis, offering weak controllability and limited support for iterative editing. Fundamentally, a world comprises multiple elements with geometry, appearance, and other attributes, together with cameras. Different frames are produced through local modifications or recombinations of this shared state, which is otherwise largely reused. Therefore, we argue that the missing component is an explicit and persistent working state. To address this, we present StateFlow, a state-centric framework for generative previsualization. Rather than generating videos in one shot, StateFlow uses an editable 3D world to organize scene structure, evolution, and cameras, while off-the-shelf video models enhance visual quality when higher fidelity is desired. This world is maintained as a persistent structured 3D state of scene elements and camera configurations, serving as the core working representation for previsualization. Built on this insight, StateFlow has three stages to construct, evolve, and access the world state. State construction lifts generated 2D content into a coherent 3D world through prior-guided, conflict-aware dual-view initialization, while State evolution translates user intent into structured state transitions while preserving world memory, avoiding full-scene regeneration for each edit. State access uses render-feedback reflection to refine camera plans into visually feasible trajectories, avoiding reliance on VLM semantics alone. Experiments show that StateFlow produces high-quality 3D worlds for video creation and game-like prototyping.