PointDiT: 픽셀 공간 확산을 통한 단안 기하 추정
PointDiT: Pixel-Space Diffusion for Monocular Geometry Estimation
July 2, 2026
저자: Haofei Xu, Rundi Wu, Philipp Henzler, Nikolai Kalischek, Michael Oechsle, Fabian Manhardt, Marc Pollefeys, Andreas Geiger, Federico Tombari, Michael Niemeyer
cs.AI
초록
최첨단 단일 이미지 3차원 재구성 방법은 종종 복잡한 하이브리드 아키텍처와 손실 함수에 의존하거나, 사전 학습된 잠재 확산 모델을 활용하기 위해 기하학 정보를 잠재 공간으로 압축한다. 본 연구에서는 이러한 아키텍처 오버헤드와 복잡한 손실 공식이 불필요함을 보여준다. 우리는 일반 ViT(비전 트랜스포머) 기반의 미니멀리스트 픽셀 공간 확산 트랜스포머를 도입하며, 이는 원시 3차원 포인트 맵 패치에 직접 작동하고 사전 학습된 DINOv3의 이미지 토큰에 의해 조건화된다. 기존의 잠재 확산 접근법과 달리, 우리는 확산 백본을 처음부터 완전히 학습시켜 포인트 맵 토크나이저의 필요성을 제거한다. 단순함에도 불구하고, 우리의 접근법은 복잡한 잠재 기반 확산 모델을 능가하면서 하이브리드 대안보다 현저히 단순하다. 특히, 더 선명한 기하학적 구조를 생성하며 투명 객체와 같이 모호성이 높은 영역에서 더 견고하다.
English
State-of-the-art single-image 3D reconstruction methods often rely on complex hybrid architectures and loss functions, or compress geometry into latent spaces in order to leverage pre-trained latent diffusion models. In this work, we show that such architectural overhead and intricate loss formulations are unnecessary. We introduce a minimalist pixel-space Diffusion Transformer, built on a plain ViT, that operates directly on raw 3D point map patches and is conditioned on image tokens from a pre-trained DINOv3. Unlike existing latent diffusion approaches, we train our diffusion backbone entirely from scratch, eliminating the need for point map tokenizers. Despite its simplicity, our approach surpasses complex latent-based diffusion models while remaining significantly simpler than hybrid alternatives. Notably, it produces sharper geometric structure and is more robust in highly ambiguous regions, such as transparent objects.