DiffGI: 고충실도 박막 3D 생성을 위한 미분 가능 기하 이미지
DiffGI: Differentiable Geometry Images for High-Fidelity Thin-Shell 3D Generation
July 15, 2026
저자: Eungjune Shim, Hansol Lee, Eunjung Ju
cs.AI
초록
기존의 3D 생성 모델은 주로 암시적 체적 표현에 의존해 왔으며, 이는 수밀 위상을 강제하고 의류와 같은 얇은 껍질 및 비다양체 기하를 표현하는 데 어려움을 겪는다. 기하 이미지 기반 접근법은 표면 중심의 대안을 제공하지만, 기존 방법은 이산 이진 점유 맵을 사용하며, 이 맵의 해상도 의존적 경계 인코딩은 다운샘플링 시 계단형 아티팩트와 정보 손실을 유발하고, 표면 재구성은 학습 파이프라인과 분리된 미분 불가능한 후처리 단계로 남아 있다. 이를 해결하기 위해, 우리는 미분 가능 기하 이미지(DiffGI)를 제안한다. 이는 표면 표현과 기하 최적화를 원활하게 통합하는 종단간 3D-2D 매핑 프레임워크이다. DiffGI는 이진 맵을 고정 격자 해상도 내에서 서브픽셀 정밀도로 경계 위치를 인코딩하는 연속 2D 절단 부호 거리 함수(TSDF)로 대체하여, 과격한 다운샘플링에서도 해상도 의존적 계단형 아티팩트를 제거한다. 이 연속 필드를 기반으로, 우리는 해석적 선형 보간에 기반한 미분 가능 마칭 스퀘어 알고리즘을 도입하여 3D 표면 손실로부터의 그래디언트가 2D 잠재 공간으로 역전파될 수 있도록 한다. 이 미분 가능 파이프라인을 활용하여, 우리는 기하 인식 법선 렌더링 손실로 강화된 DiffGI-VAE를 훈련시켜 복잡한 3D 표면을 초소형 32X32 잠재 공간으로 압축하고, 이 잠재 공간 위에 흐름 매칭 목적 함수를 가진 트랜스포머 기반 잠재 확산 모델을 인스턴스화하여 조건부 3D 생성을 수행한다. 의상 및 객체 데이터셋에 대한 광범위한 실험은 우리의 방법이 이전의 기하 이미지 및 복셀 기반 접근법에 비해 우수한 재구성 충실도와 경계 정밀도를 달성하면서도 훨씬 적은 계산 자원을 필요로 함을 보여준다.
English
Existing 3D generative models predominantly rely on implicit volumetric representations, which enforce watertight topology and struggle to represent thin-shell and non-manifold geometries such as garments. Geometry image-based approaches offer a surface-centric alternative, but existing methods rely on discrete binary occupancy maps whose resolution-dependent boundary encoding causes staircase artifacts and information loss upon downsampling, while surface reconstruction remains a non-differentiable post-processing step disconnected from the learning pipeline. To address this, we propose Differentiable Geometry Image (DiffGI), an end-to-end 3D-to-2D mapping framework that seamlessly integrates surface representation and geometric optimization. DiffGI replaces binary maps with a continuous 2D Truncated Signed Distance Function (TSDF), which encodes boundary position at subpixel precision within a fixed grid resolution, eliminating resolution-dependent staircase artifacts even under aggressive downsampling. Building on this continuous field, we introduce a differentiable Marching Squares algorithm based on analytical linear interpolation, allowing gradients from 3D surface losses to propagate back to the 2D latent space. Leveraging this differentiable pipeline, we train a DiffGI-VAE augmented with a geometry-aware normal rendering loss to compress complex 3D surfaces into an ultra-compact 32X32 latent space, and instantiate a transformer-based latent diffusion model with a flow-matching objective on top of this space for conditional 3D generation. Extensive experiments on garment and object datasets demonstrate that our method achieves superior reconstruction fidelity and boundary precision compared to prior geometry-image and voxel-based approaches, while requiring significantly fewer computational resources.