ChatPaper.aiChatPaper

GS-Voxel:用於大規模3DGS生成的免擬合結構化潛在變量

GS-Voxel: Fitting-Free Structured Latents for Large-Scale 3DGS Generation

August 18, 2026
作者: Ming Qian, Zijian Wang, Minchao Sun, Jincheng Xiong, Hang Zhang, Mu Xu, Chi Wang, Baoquan Chen
cs.AI

摘要

許多可擴展的潛在 3D 生成器作用於結構化張量,而預先優化的 3D 高斯潑濺(3DGS)重建則是無序、空間上不規則,且圖元數量差異極大。我們提出 GS-Voxel,一個免擬合的結構化潛在框架,並對其進行大規模空中 3D 高斯場景生成的評估。GS-Voxel 可將相容的預先優化 3DGS 重建確定性轉換為稀疏活躍體素,無需額外的逐場景優化,同時保留所選圖元的亞體素位置與渲染屬性。接著,一個 GS 特定的分解 VAE 將體素幾何與局部高斯屬性分別編碼為稀疏 3D 潛在表示,其大小隨佔用體素數量增長,而非受限於固定的場景範圍圖元數量。我們在 GS-Voxel 潛在空間中訓練以影像為條件的流模型,以生成空中 3DGS 場景。GS-Voxel 所實現的一項關鍵應用是大面積場景生成:重疊感知的分塊推論可將合成範圍擴展至單一訓練裁切之外,並以衛星視圖影像為條件。我們的結果顯示,GS-Voxel 為預先優化的空中 3DGS 重建提供了結構化潛在表示,且其潛在容量隨佔用體素數量增長。
English
Many scalable latent 3D generators operate on structured tensors, whereas pre-optimized 3D Gaussian Splatting (3DGS) reconstructions are unordered, spatially irregular, and vary widely in primitive count. We present GS-Voxel, a fitting-free structured latent framework, and evaluate it for large-scale aerial 3D Gaussian scene generation. GS-Voxel deterministically converts a compatible pre-optimized 3DGS reconstruction into sparse active voxels without additional per-scene optimization, retaining the sub-voxel positions and rendering attributes of the selected primitives. A GS-specific factorized VAE then separately encodes voxel geometry and local Gaussian attributes into sparse 3D latents whose size grows with the number of occupied voxels rather than being limited by a fixed scene-wide primitive count. We train image-conditioned flow models in the GS-Voxel latent space to generate aerial 3DGS scenes. A key application enabled by GS-Voxel is large-area scene generation: overlap-aware tiled inference extends synthesis beyond a single training crop conditioned on satellite-view images. Our results show that GS-Voxel provides structured latents for pre-optimized aerial 3DGS reconstructions, with latent capacity that grows with the number of occupied voxels.