ChatPaper.aiChatPaper

PixWorld: ピクセル空間における3Dシーン生成と再構成の統合

PixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Space

July 6, 2026
著者: Sensen Gao, Zhaoqing Wang, Qihang Cao, Dongdong Yu, Changhu Wang, Jia-Wang Bian
cs.AI

要旨

3次元再構築と生成は通常、別々のパラダイム、すなわち再構築には画素ベースの回帰、生成には潜在拡散モデルによって扱われてきた。近年の研究ではこれらを潜在空間で統合しようと試みているが、顕著な欠点がある。拡散目的関数が基礎となる3次元表現ではなく潜在特徴量に定義されていること、そして両方のブランチが潜在エンコーディングによる情報損失の影響を受け、さらに事前学習された変分オートエンコーダ(VAE)または表現オートエンコーダ(RAE)を必要とすることである。本論文では、これら二つのタスクを統一された画素空間拡散パラダイムの下で再定式化し、3次元再構築と生成を同時に扱う単一モデルPixWorldを提案する。拡散をレンダリング画像に直接適用することで、PixWorldは上記の制約を取り除き、最適化を3次元シーンの忠実性と整合させる。2次元画像レベルで動作し3次元幾何学的認識を欠く光度および知覚的監督に加えて、さらに幾何認識損失を導入する。これは、事前学習された3次元基盤モデルの幾何認識特徴空間において、レンダリングされたビューをその正解データと整合させ、3次元構造的監督を提供する。PixWorldは、従来の潜在空間生成手法を一貫して上回り、最先端の再構築手法と同等の性能を示し、統合された画素空間アプローチの優位性を実証している。
English
3D reconstruction and generation are commonly tackled by separate paradigms: pixel-based regression for reconstruction, and latent diffusion for generation. Recent works attempt to unify them in latent space, but with notable drawbacks: the diffusion objective is defined on latent features rather than the underlying 3D representation, and both branches suffer from information loss introduced by latent encoding, while requiring a pretrained Variational Autoencoder (VAE) or Representation Autoencoder (RAE). In this paper, we reformulate these two tasks under a unified pixel-space diffusion paradigm and introduce PixWorld, a single model that jointly addresses 3D reconstruction and generation. By supervising diffusion directly on rendered images, PixWorld removes the above limitations and aligns optimization with 3D scene fidelity. Beyond photometric and perceptual supervision that operates at the 2D image level and lacks 3D geometric awareness, we further introduce a geometry perception loss that aligns rendered views with their ground truth in the geometry-aware feature space of a pretrained 3D foundation model, providing 3D structural supervision. PixWorld consistently outperforms prior latent-space generation methods and matches state-of-the-art reconstruction methods, demonstrating the superiority of a unified pixel-space approach.