ChatPaper.aiChatPaper

SynCity 3000: 장면 규모 3D 확산의 부트스트래핑

SynCity 3000: Bootstrapping Scene-Scale 3D Diffusion

July 6, 2026
저자: Paul Engstler, Iro Laina, Christian Rupprecht, Andrea Vedaldi
cs.AI

초록

SynCity 3000을 제시한다. 이 프레임워크는 전역적으로 일관된 3D 장면을 생성하면서도 세밀한 배치 제어를 가능하게 한다. 단일 이미지로부터 복잡한 3D 자산을 생성할 수 있는 현재의 이미지-3D 생성기의 능력을 바탕으로, 생성기를 합성곱 연산자로 적용할 수 있도록 조정하여 이러한 능력을 전체 장면 규모로 확장한다. 이를 위해 훈련용 3D 장면 데이터의 부족 문제를 해결하기 위해 제안하는 새로운 합성 데이터 엔진으로 생성된 장면 유사 데이터에 모델을 미세 조정한다. 그런 다음 합성곱 생성기를 사용자 프롬프트로부터 생성된 전체 장면의 이측면 이미지에 적용하여 임의의 크기와 복잡성을 가진 3D 장면을 생성한다. 다양한 프롬프트와 배치에 걸쳐 SynCity 3000은 크고 일관되며 세부적인 장면을 생성하여 기존 3D 장면 생성 접근 방식의 한계를 해결한다.
English
We present SynCity 3000, a framework for generating 3D scenes that are globally coherent while enabling fine-grained layout control. Building on the ability of current image-to-3D generators to produce complex 3D assets from a single image, we extend this capability to the scale of entire scenes by adapting the generator to be applicable as a convolutional operator. We achieve this by fine-tuning the model on scene-like data generated by a new synthetic data engine, which we propose to address the scarcity of 3D scene data for training. The convolutional generator is then applied to a dimetric image of the entire scene, generated from the user prompt, resulting in 3D scenes of arbitrary size and complexity. Across diverse prompts and layouts, SynCity 3000 produces large, coherent, and detailed scenes, addressing the shortcomings of prior approaches to 3D scene generation.