ATSplat: 適応的トークン拡張によるコンパクトなフィードフォワード3Dガウシアンスプラッティング
ATSplat: Compact Feed-forward 3D Gaussian Splatting with Adaptive Token Expansion
July 22, 2026
著者: Cho In, Jeonghwan Cho, Mijin Yoo, Gim Hee Lee, Seon Joo Kim
cs.AI
要旨
3D Gaussian Splatting(3DGS)は、3次元空間に自由に配置されたプリミティブを最適化し、未再構成領域では適応的に高密度化することで、高品質な新規視点合成を実現する。しかし、既存のフィードフォワード型3DGS手法では、このシーン適応的な容量割り当てがほとんど失われており、一般的に入力ピクセル上でガウシアンを回帰し、それらをカメラ光線に沿って持ち上げている。このようなピクセル整列型の定式化では、プリミティブの数と配置がシーンの複雑さではなく画像解像度や入力視点に依存するため、結果として密集した冗長なガウシアン集合となる。本稿では、適応型3Dトークン(Adaptive 3D Tokens)を通じて3DGS最適化の適応的割り当て能力を回復する、フィードフォワード型3DGSフレームワークであるATSplatを提案する。ATSplatはまず、粗いパッチレベルの深度とカメラ情報を疎な3Dアンカートークンに持ち上げ、シーンのコンパクトな足場を形成する。各トークンは、学習可能な3Dオフセットを用いて局所的なガウシアンに回帰され、プリミティブの配置を入力画像グリッドから切り離す。適応型トークン拡張(Adaptive Token Expansion)モジュールは、レンダリング誤差マップで教師されるトークンレベルの不確実性スコアを予測し、学習可能な拡張層を通じて高不確実性トークンを選択的に拡張する。このスパースから適応的への定式化により、ATSplatはチャレンジングな領域にプリミティブを集中させつつ、コンパクトな表現を維持できる。RealEstate10KとDL3DVの2つの代表的なデータセットにおける実験では、ATSplatは高密度なフィードフォワード型3DGS手法と比較してガウシアン数を5.7倍以上削減しつつ、最先端のレンダリング品質を達成する。512×960の解像度の12枚の入力画像から、ATSplatは1枚の市販GPUを用いて1秒未満で再構成を完了し、わずか311Kのガウシアンで1136 FPS(512×960)の高品質な新規視点レンダリングを実現する。
English
3D Gaussian Splatting (3DGS) achieves high-quality novel-view synthesis by optimizing freely placed primitives in 3D and adaptively densifying them in under-reconstructed regions. However, this scene-adaptive capacity allocation is largely lost in existing feed-forward 3DGS methods, which commonly regress Gaussians at input pixels and lift them along camera rays. Such pixel-aligned formulations make the number and placement of primitives depend on image resolution and input viewpoints rather than scene complexity, resulting in dense and often redundant Gaussian sets. We present ATSplat, a feed-forward 3DGS framework that restores the adaptive allocation capability of 3DGS optimization through Adaptive 3D Tokens. ATSplat first lifts coarse patch-level depth and camera cues into sparse 3D anchor tokens, forming a compact scaffold of the scene. Each token is then regressed into local Gaussians with learnable 3D offsets, decoupling primitive placement from input image grids. An Adaptive Token Expansion module predicts a token-level uncertainty score, supervised by rendering error maps, and selectively expands high-uncertainty tokens through learnable expansion layers. This sparse-to-adaptive formulation enables ATSplat to concentrate primitives in challenging regions while maintaining a compact representation. Experiments on two representative datasets, RealEstate10K and DL3DV, show that ATSplat achieves state-of-the-art rendering quality while reducing the number of Gaussians by more than 5.7times compared with dense feed-forward 3DGS methods. From 12 input images at 512 times 960 resolution, ATSplat completes reconstruction in less than a second using a single commercial GPU, and renders high-quality novel views at 1136 FPS (512 times 960) with only 311K Gaussians.