AsySplat: 効率的な非対称3Dガウシアンスプラッティングによる長シーケンスシーンモデリング
AsySplat: Efficient Asymmetric 3D Gaussian Splatting for Long-Sequence Scene Modeling
July 13, 2026
著者: Yingji Zhong, Dave Zhenyu Chen, Fuzhao Ou, Youyu Chen, Zhihao Li, Lanqing Hong, Dan Xu
cs.AI
要旨
近年の汎化可能な3Dガウシアンスプラッティングモデルは長系列の新規視点合成(NVS)を進展させたが、その代償として大量の冗長な計算を伴う。我々は、この冗長性が2つの観察に基づいて軽減可能であることを特定する:(i) 高品質なNVSには高精度な幾何学は厳密には必要ないこと、(ii) 外観学習は一般的に幾何学復元よりも容易であること。これらの洞察に動機づけられ、我々は幾何学と外観モデリングを分離する非対称アーキテクチャを提案する。幾何学ブランチは粗粒度トークンを処理し、多視点再構成のためにほとんどのパラメータを使用する一方、外観ブランチは細粒度トークンを操作し、大幅に少ないパラメータで詳細を捉える。2つのブランチは双方向接続を通じて相互作用し、それぞれのタスクの相互ガイダンスを可能にする。このタスク認識非対称性は計算の冗長性を低減し、計算をより適切に配分することで、パラメータ効率を高め、より小規模なモデルでも強力な性能を達成できるようにする。32視点の960P入力において、我々のモデルは最適化ベースの手法に匹敵しつつ、約800倍の高速化を実現し、顕著に少ないパラメータと低減された訓練/推論オーバーヘッドで最先端の汎化可能モデルのゼロショット性能を上回り、全体的な効率改善を達成する。
English
Recent generalizable 3D Gaussian Splatting models have advanced long-sequence novel view synthesis (NVS), but at the cost of substantial redundant computation. We identify that the redundancy can be mitigated based on two observations: (i) high-precision geometry is not strictly required for high-quality NVS; (ii) appearance learning is generally easier than geometry recovery. Motivated by these insights, we propose an asymmetric architecture that decouples geometry and appearance modeling. The geometry branch processes coarse-grained tokens with most of the parameters for multi-view reconstruction, while the appearance branch operates on fine-grained tokens to capture details using significantly fewer parameters. The two branches interact through bilateral connections, enabling mutual guidance for their respective tasks. This task-aware asymmetry reduces the computational redundancy and allocates the computation more judiciously, thereby increasing parameter efficiency and enabling smaller models to achieve strong performance. On 32-view 960P inputs, our model matches optimization-based methods while delivering nearly 800x speedup, and surpasses the zero-shot performance of state-of-the-art generalizable models with markedly fewer parameters and reduced training/inference overhead, achieving an overall efficiency improvement.