ChatPaper.aiChatPaper

SynCity 3000:自舉場景級三維擴散

SynCity 3000: Bootstrapping Scene-Scale 3D Diffusion

July 6, 2026
作者: Paul Engstler, Iro Laina, Christian Rupprecht, Andrea Vedaldi
cs.AI

摘要

我們提出 SynCity 3000,這是一個用於生成全局一致、同時支援細粒度佈局控制的 3D 場景框架。奠基於當前圖像轉 3D 生成器從單張圖像中產生複雜 3D 資產的能力,我們透過將生成器改造為可當作卷積運算元來應用,將此能力擴展至整個場景的規模。我們透過在場景類資料上微調模型來實現此目標,這些資料由我們提出的新型合成資料引擎所生成,旨在解決訓練中 3D 場景資料稀缺的問題。接著,卷積生成器被應用於根據使用者提示所生成的整個場景之斜二測圖像,從而產生任意大小與複雜度的 3D 場景。在各種提示和佈局中,SynCity 3000 生成了大型、連貫且細節豐富的場景,解決了先前 3D 場景生成方法的不足。
English
We present SynCity 3000, a framework for generating 3D scenes that are globally coherent while enabling fine-grained layout control. Building on the ability of current image-to-3D generators to produce complex 3D assets from a single image, we extend this capability to the scale of entire scenes by adapting the generator to be applicable as a convolutional operator. We achieve this by fine-tuning the model on scene-like data generated by a new synthetic data engine, which we propose to address the scarcity of 3D scene data for training. The convolutional generator is then applied to a dimetric image of the entire scene, generated from the user prompt, resulting in 3D scenes of arbitrary size and complexity. Across diverse prompts and layouts, SynCity 3000 produces large, coherent, and detailed scenes, addressing the shortcomings of prior approaches to 3D scene generation.