ChatPaper.aiChatPaper

PointDiT: ピクセル空間拡散による単眼幾何推定

PointDiT: Pixel-Space Diffusion for Monocular Geometry Estimation

July 2, 2026
著者: Haofei Xu, Rundi Wu, Philipp Henzler, Nikolai Kalischek, Michael Oechsle, Fabian Manhardt, Marc Pollefeys, Andreas Geiger, Federico Tombari, Michael Niemeyer
cs.AI

要旨

最先端の単一画像からの3次元再構成手法は、しばしば複雑なハイブリッドアーキテクチャや損失関数に依存するか、または事前学習済みの潜在拡散モデルを活用するために幾何形状を潜在空間に圧縮します。本研究では、このような構造的なオーバーヘッドや複雑な損失関数の定式化は不要であることを示します。我々は、プレーンなViT上に構築されたミニマルなピクセル空間の拡散トランスフォーマーを導入します。これは、生の3次元点群マップパッチに直接作用し、事前学習済みのDINOv3からの画像トークンによって条件付けられます。既存の潜在拡散手法とは異なり、我々は拡散バックボーンを完全にスクラッチから学習し、点群マップ用のトークナイザーを不要にします。その単純さにもかかわらず、我々の手法は複雑な潜在ベースの拡散モデルを凌駕し、ハイブリッド手法よりもはるかに単純なままです。特に、よりシャープな幾何学的構造を生成し、透明な物体などの非常に曖昧な領域においてより頑健です。
English
State-of-the-art single-image 3D reconstruction methods often rely on complex hybrid architectures and loss functions, or compress geometry into latent spaces in order to leverage pre-trained latent diffusion models. In this work, we show that such architectural overhead and intricate loss formulations are unnecessary. We introduce a minimalist pixel-space Diffusion Transformer, built on a plain ViT, that operates directly on raw 3D point map patches and is conditioned on image tokens from a pre-trained DINOv3. Unlike existing latent diffusion approaches, we train our diffusion backbone entirely from scratch, eliminating the need for point map tokenizers. Despite its simplicity, our approach surpasses complex latent-based diffusion models while remaining significantly simpler than hybrid alternatives. Notably, it produces sharper geometric structure and is more robust in highly ambiguous regions, such as transparent objects.