PixWorld: 픽셀 공간에서 3D 장면 생성 및 재구성 통합
PixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Space
July 6, 2026
저자: Sensen Gao, Zhaoqing Wang, Qihang Cao, Dongdong Yu, Changhu Wang, Jia-Wang Bian
cs.AI
초록
3D 재구성과 생성은 일반적으로 픽셀 기반 회귀를 통한 재구성과 잠재 확산을 통한 생성이라는 별도의 패러다임으로 다루어져 왔다. 최근 연구들은 이를 잠재 공간에서 통합하려 시도했지만, 확산 목적 함수가 기본 3D 표현이 아닌 잠재 특징에 대해 정의되고, 두 분기 모두 잠재 인코딩으로 인한 정보 손실을 겪으며 사전 훈련된 변분 오토인코더(VAE) 또는 표현 오토인코더(RAE)를 필요로 한다는 뚜렷한 단점이 있다. 본 논문에서는 이 두 작업을 통합된 픽셀 공간 확산 패러다임으로 재정립하고, 3D 재구성과 생성 문제를 공동으로 해결하는 단일 모델인 PixWorld를 제안한다. 확산 과정을 렌더링된 이미지에 직접 감독함으로써 PixWorld는 위의 한계를 제거하고 최적화를 3D 장면 충실도와 정렬시킨다. 2D 이미지 수준에서 작동하여 3D 기하학적 인식이 부족한 측광 및 지각적 감독을 넘어, 사전 훈련된 3D 기반 모델의 기하학 인식 특징 공간에서 렌더링된 뷰를 실제 값과 정렬하는 기하학 지각 손실을 추가로 도입하여 3D 구조적 감독을 제공한다. PixWorld는 기존의 잠재 공간 생성 방법보다 일관되게 우수한 성능을 보이며 최첨단 재구성 방법과 동등한 수준을 달성함으로써 통합된 픽셀 공간 접근 방식의 우월성을 입증한다.
English
3D reconstruction and generation are commonly tackled by separate paradigms: pixel-based regression for reconstruction, and latent diffusion for generation. Recent works attempt to unify them in latent space, but with notable drawbacks: the diffusion objective is defined on latent features rather than the underlying 3D representation, and both branches suffer from information loss introduced by latent encoding, while requiring a pretrained Variational Autoencoder (VAE) or Representation Autoencoder (RAE). In this paper, we reformulate these two tasks under a unified pixel-space diffusion paradigm and introduce PixWorld, a single model that jointly addresses 3D reconstruction and generation. By supervising diffusion directly on rendered images, PixWorld removes the above limitations and aligns optimization with 3D scene fidelity. Beyond photometric and perceptual supervision that operates at the 2D image level and lacks 3D geometric awareness, we further introduce a geometry perception loss that aligns rendered views with their ground truth in the geometry-aware feature space of a pretrained 3D foundation model, providing 3D structural supervision. PixWorld consistently outperforms prior latent-space generation methods and matches state-of-the-art reconstruction methods, demonstrating the superiority of a unified pixel-space approach.