ChatPaper.aiChatPaper

SynCity 3000:シーンスケール3D拡散のブートストラッピング

SynCity 3000: Bootstrapping Scene-Scale 3D Diffusion

July 6, 2026
著者: Paul Engstler, Iro Laina, Christian Rupprecht, Andrea Vedaldi
cs.AI

要旨

SynCity 3000を提案する。これは、全体として一貫性を持ちながら、きめ細かいレイアウト制御を可能にする3Dシーン生成フレームワークである。現在の画像から3Dへの生成器が単一画像から複雑な3Dアセットを生成できる能力を基盤とし、この生成器を畳み込み演算子として適用可能に適応させることで、その能力をシーン全体の規模に拡張する。これは、モデルを新たな合成データエンジンによって生成されたシーン状データでファインチューニングすることで実現する。このデータエンジンは、トレーニング用の3Dシーンデータの不足に対処するために提案する。次に、畳み込み生成器を、ユーザープロンプトから生成されたシーン全体の二軸測投影画像に適用し、任意のサイズと複雑さを持つ3Dシーンを生成する。多様なプロンプトとレイアウトにわたって、SynCity 3000は大規模で一貫性があり、かつ詳細なシーンを生成し、従来の3Dシーン生成手法の欠点に対処する。
English
We present SynCity 3000, a framework for generating 3D scenes that are globally coherent while enabling fine-grained layout control. Building on the ability of current image-to-3D generators to produce complex 3D assets from a single image, we extend this capability to the scale of entire scenes by adapting the generator to be applicable as a convolutional operator. We achieve this by fine-tuning the model on scene-like data generated by a new synthetic data engine, which we propose to address the scarcity of 3D scene data for training. The convolutional generator is then applied to a dimetric image of the entire scene, generated from the user prompt, resulting in 3D scenes of arbitrary size and complexity. Across diverse prompts and layouts, SynCity 3000 produces large, coherent, and detailed scenes, addressing the shortcomings of prior approaches to 3D scene generation.