ChatPaper.aiChatPaper

GS-Voxel: 대규모 3DGS 생성을 위한 피팅-프리 구조화된 잠재 표현

GS-Voxel: Fitting-Free Structured Latents for Large-Scale 3DGS Generation

August 18, 2026
저자: Ming Qian, Zijian Wang, Minchao Sun, Jincheng Xiong, Hang Zhang, Mu Xu, Chi Wang, Baoquan Chen
cs.AI

초록

확장 가능한 잠재(latent) 3D 생성기 다수는 구조화된 텐서를 기반으로 동작하는 반면, 사전 최적화된 3D 가우시안 스플래팅(3DGS) 재구성은 비정렬적이고 공간적으로 불규칙하며 프리미티브 개수도 매우 다양하다. 본 논문에서는 피팅(fitting)이 필요 없는 구조화된 잠재 프레임워크인 GS-Voxel을 제시하고, 대규모 항공 3D 가우시안 장면 생성을 위해 이를 평가한다. GS-Voxel은 호환 가능한 사전 최적화된 3DGS 재구성을 추가적인 장면별 최적화 없이 결정론적으로 희소 활성 복셀로 변환하며, 선택된 프리미티브의 서브복셀 위치와 렌더링 속성을 유지한다. 이후 GS 특화 분해(factorized) VAE가 복셀 형상과 로컬 가우시안 속성을 별도로 희소 3D 잠재 표현으로 인코딩하는데, 이 잠재 표현의 크기는 고정된 장면 전체 프리미티브 개수에 의해 제한되지 않고 점유된 복셀의 수에 따라 증가한다. 또한 GS-Voxel 잠재 공간에서 이미지 조건부 플로우 모델을 학습하여 항공 3DGS 장면을 생성한다. GS-Voxel이 가능하게 하는 핵심 응용은 대면적 장면 생성으로, 중첩 인식 타일 추론(overlap-aware tiled inference)은 위성 뷰 이미지에 조건화된 단일 학습 크롭을 넘어 합성을 확장한다. 실험 결과는 GS-Voxel이 사전 최적화된 항공 3DGS 재구성을 위한 구조화된 잠재 표현을 제공하며, 잠재 용량이 점유된 복셀 수에 따라 증가함을 보여준다.
English
Many scalable latent 3D generators operate on structured tensors, whereas pre-optimized 3D Gaussian Splatting (3DGS) reconstructions are unordered, spatially irregular, and vary widely in primitive count. We present GS-Voxel, a fitting-free structured latent framework, and evaluate it for large-scale aerial 3D Gaussian scene generation. GS-Voxel deterministically converts a compatible pre-optimized 3DGS reconstruction into sparse active voxels without additional per-scene optimization, retaining the sub-voxel positions and rendering attributes of the selected primitives. A GS-specific factorized VAE then separately encodes voxel geometry and local Gaussian attributes into sparse 3D latents whose size grows with the number of occupied voxels rather than being limited by a fixed scene-wide primitive count. We train image-conditioned flow models in the GS-Voxel latent space to generate aerial 3DGS scenes. A key application enabled by GS-Voxel is large-area scene generation: overlap-aware tiled inference extends synthesis beyond a single training crop conditioned on satellite-view images. Our results show that GS-Voxel provides structured latents for pre-optimized aerial 3DGS reconstructions, with latent capacity that grows with the number of occupied voxels.