PixWorld:在像素空間中統一3D場景生成與重建
PixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Space
July 6, 2026
作者: Sensen Gao, Zhaoqing Wang, Qihang Cao, Dongdong Yu, Changhu Wang, Jia-Wang Bian
cs.AI
摘要
3D重建與生成通常由不同的範疇處理:基於像素的回歸用於重建,潛在擴散用於生成。近期研究嘗試在潛在空間中統一兩者,但存在明顯缺點:擴散目標定義於潛在特徵而非底層3D表示,且兩個分支均受潛在編碼引入的資訊損失影響,同時需依賴預訓練的變分自動編碼器(VAE)或表示自動編碼器(RAE)。本文將這兩項任務重新架構於統一的像素空間擴散範疇,提出PixWorld——一個同時處理3D重建與生成的單一模型。透過直接在渲染影像上監督擴散,PixWorld消除了上述限制,並使優化與3D場景保真度對齊。除了在2D影像層級運作且缺乏3D幾何感知的光度與感知監督外,我們進一步引進幾何感知損失,將渲染視圖與其在預訓練3D基礎模型的幾何感知特徵空間中的真實標註對齊,提供3D結構監督。PixWorld持續優於先前的潛在空間生成方法,並可與最先進的重建方法匹敵,展現了統一像素空間方法的優越性。
English
3D reconstruction and generation are commonly tackled by separate paradigms: pixel-based regression for reconstruction, and latent diffusion for generation. Recent works attempt to unify them in latent space, but with notable drawbacks: the diffusion objective is defined on latent features rather than the underlying 3D representation, and both branches suffer from information loss introduced by latent encoding, while requiring a pretrained Variational Autoencoder (VAE) or Representation Autoencoder (RAE). In this paper, we reformulate these two tasks under a unified pixel-space diffusion paradigm and introduce PixWorld, a single model that jointly addresses 3D reconstruction and generation. By supervising diffusion directly on rendered images, PixWorld removes the above limitations and aligns optimization with 3D scene fidelity. Beyond photometric and perceptual supervision that operates at the 2D image level and lacks 3D geometric awareness, we further introduce a geometry perception loss that aligns rendered views with their ground truth in the geometry-aware feature space of a pretrained 3D foundation model, providing 3D structural supervision. PixWorld consistently outperforms prior latent-space generation methods and matches state-of-the-art reconstruction methods, demonstrating the superiority of a unified pixel-space approach.