GS-Voxel:面向大规模3DGS生成的免拟合结构化潜在表示
GS-Voxel: Fitting-Free Structured Latents for Large-Scale 3DGS Generation
August 18, 2026
作者: Ming Qian, Zijian Wang, Minchao Sun, Jincheng Xiong, Hang Zhang, Mu Xu, Chi Wang, Baoquan Chen
cs.AI
摘要
许多可扩展的潜在3D生成器基于结构化张量进行操作,而预优化的3D高斯溅射(3DGS)重建则是无序、空间不规则且基元数量差异很大的。我们提出了GS-Voxel,一种免拟合的结构化潜在框架,并针对大规模航空3D高斯场景生成进行了评估。GS-Voxel将兼容的预优化3DGS重建确定性转换为稀疏活动体素,无需额外的逐场景优化,同时保留所选基元的亚体素位置和渲染属性。随后,一个针对GS的分解VAE将体素几何和局部高斯属性分别编码为稀疏3D潜在表示,其大小随占据体素数量增长,而非受限于固定的场景级基元数量。我们在GS-Voxel潜在空间中训练图像条件流模型,以生成航空3DGS场景。GS-Voxel实现的一个关键应用是大面积场景生成:重叠感知的分块推理将基于卫星视角图像的合成扩展到单个训练裁剪之外。我们的结果表明,GS-Voxel为预优化的航空3DGS重建提供了结构化潜在表示,其潜在容量随占据体素数量增长。
English
Many scalable latent 3D generators operate on structured tensors, whereas pre-optimized 3D Gaussian Splatting (3DGS) reconstructions are unordered, spatially irregular, and vary widely in primitive count. We present GS-Voxel, a fitting-free structured latent framework, and evaluate it for large-scale aerial 3D Gaussian scene generation. GS-Voxel deterministically converts a compatible pre-optimized 3DGS reconstruction into sparse active voxels without additional per-scene optimization, retaining the sub-voxel positions and rendering attributes of the selected primitives. A GS-specific factorized VAE then separately encodes voxel geometry and local Gaussian attributes into sparse 3D latents whose size grows with the number of occupied voxels rather than being limited by a fixed scene-wide primitive count. We train image-conditioned flow models in the GS-Voxel latent space to generate aerial 3DGS scenes. A key application enabled by GS-Voxel is large-area scene generation: overlap-aware tiled inference extends synthesis beyond a single training crop conditioned on satellite-view images. Our results show that GS-Voxel provides structured latents for pre-optimized aerial 3DGS reconstructions, with latent capacity that grows with the number of occupied voxels.