GS-Voxel: フィッティング不要の構造化潜在表現による大規模3DGS生成
GS-Voxel: Fitting-Free Structured Latents for Large-Scale 3DGS Generation
August 18, 2026
著者: Ming Qian, Zijian Wang, Minchao Sun, Jincheng Xiong, Hang Zhang, Mu Xu, Chi Wang, Baoquan Chen
cs.AI
要旨
多くのスケーラブルな潜在3D生成モデルは構造化テンソル上で動作するのに対し、事前最適化された3Dガウススプラッティング(3DGS)再構成は非順序的で空間的に不規則であり、プリミティブ数も大きく異なる。我々は、フィッティング不要の構造化潜在フレームワークであるGS-Voxelを提案し、大規模な航空3Dガウスシーン生成におけるその有効性を評価する。GS-Voxelは、互換性のある事前最適化3DGS再構成を、シーンごとの追加最適化なしで、選択されたプリミティブのサブボクセル位置とレンダリング属性を保持したまま、スパースなアクティブボクセルへ決定論的に変換する。次に、GS特化の因子分解VAEが、ボクセルジオメトリと局所的なガウス属性を別々にスパースな3D潜在表現へ符号化する。この潜在表現のサイズは、シーン全体の固定プリミティブ数に制限されるのではなく、占有ボクセル数とともに増大する。我々は、航空3DGSシーンを生成するために、GS-Voxel潜在空間において画像条件付きフローモデルを訓練する。GS-Voxelによって可能になる重要な応用は、大面積シーン生成である。すなわち、重複を考慮したタイル推論は、衛星視点画像を条件として、単一のトレーニングクロップを超えて合成を拡張する。我々の結果は、GS-Voxelが事前最適化された航空3DGS再構成に対して構造化潜在表現を提供し、その潜在容量が占有ボクセル数とともに増大することを示している。
English
Many scalable latent 3D generators operate on structured tensors, whereas pre-optimized 3D Gaussian Splatting (3DGS) reconstructions are unordered, spatially irregular, and vary widely in primitive count. We present GS-Voxel, a fitting-free structured latent framework, and evaluate it for large-scale aerial 3D Gaussian scene generation. GS-Voxel deterministically converts a compatible pre-optimized 3DGS reconstruction into sparse active voxels without additional per-scene optimization, retaining the sub-voxel positions and rendering attributes of the selected primitives. A GS-specific factorized VAE then separately encodes voxel geometry and local Gaussian attributes into sparse 3D latents whose size grows with the number of occupied voxels rather than being limited by a fixed scene-wide primitive count. We train image-conditioned flow models in the GS-Voxel latent space to generate aerial 3DGS scenes. A key application enabled by GS-Voxel is large-area scene generation: overlap-aware tiled inference extends synthesis beyond a single training crop conditioned on satellite-view images. Our results show that GS-Voxel provides structured latents for pre-optimized aerial 3DGS reconstructions, with latent capacity that grows with the number of occupied voxels.