ChatPaper.aiChatPaper

ATSplat: 적응형 토큰 확장을 이용한 컴팩트 피드포워드 3D 가우시안 스플래팅

ATSplat: Compact Feed-forward 3D Gaussian Splatting with Adaptive Token Expansion

July 22, 2026
저자: Cho In, Jeonghwan Cho, Mijin Yoo, Gim Hee Lee, Seon Joo Kim
cs.AI

초록

3D 가우시안 스플래팅(3DGS)은 3D 공간에서 자유롭게 배치된 프리미티브를 최적화하고 재구성이 부족한 영역에서 적응적으로 밀도화함으로써 고품질의 새로운 시점 합성을 달성한다. 그러나 기존 피드포워드 3DGS 방법에서는 이러한 장면 적응형 용량 할당이 대부분 상실되는데, 이들 방법은 일반적으로 입력 픽셀에서 가우시안을 회귀하고 이를 카메라 광선을 따라 리프트한다. 이러한 픽셀 정렬 방식은 프리미티브의 수와 배치를 장면 복잡도가 아닌 이미지 해상도와 입력 시점에 의존하게 하여, 밀집되고 종종 중복되는 가우시안 집합을 초래한다. 본 논문에서는 적응형 3D 토큰(Adaptive 3D Tokens)을 통해 3DGS 최적화의 적응형 할당 능력을 복원하는 피드포워드 3DGS 프레임워크인 ATSplat을 제안한다. ATSplat은 먼저 거친 패치 수준의 깊이와 카메라 정보를 희소한 3D 앵커 토큰으로 리프트하여 장면의 간결한 스캐폴드를 형성한다. 그런 다음 각 토큰을 학습 가능한 3D 오프셋을 사용하여 로컬 가우시안으로 회귀하여, 프리미티브 배치를 입력 이미지 그리드로부터 분리한다. 적응형 토큰 확장(Adaptive Token Expansion) 모듈은 렌더링 오류 맵으로 감독되는 토큰 수준의 불확실성 점수를 예측하고, 학습 가능한 확장 레이어를 통해 불확실성이 높은 토큰을 선택적으로 확장한다. 이러한 희소-적응형 구성은 ATSplat이 간결한 표현을 유지하면서 까다로운 영역에 프리미티브를 집중시킬 수 있게 한다. 두 개의 대표적인 데이터셋인 RealEstate10K와 DL3DV에 대한 실험 결과, ATSplat은 밀집 피드포워드 3DGS 방법들과 비교하여 가우시안 수를 5.7배 이상 줄이면서 최신 수준의 렌더링 품질을 달성함을 보여준다. 512×960 해상도의 12개의 입력 이미지로부터 ATSplat은 단일 상용 GPU를 사용하여 1초 미만으로 재구성을 완료하고, 단 311K개의 가우시안만으로 1136 FPS(512×960 해상도)의 고품질 새로운 시점을 렌더링한다.
English
3D Gaussian Splatting (3DGS) achieves high-quality novel-view synthesis by optimizing freely placed primitives in 3D and adaptively densifying them in under-reconstructed regions. However, this scene-adaptive capacity allocation is largely lost in existing feed-forward 3DGS methods, which commonly regress Gaussians at input pixels and lift them along camera rays. Such pixel-aligned formulations make the number and placement of primitives depend on image resolution and input viewpoints rather than scene complexity, resulting in dense and often redundant Gaussian sets. We present ATSplat, a feed-forward 3DGS framework that restores the adaptive allocation capability of 3DGS optimization through Adaptive 3D Tokens. ATSplat first lifts coarse patch-level depth and camera cues into sparse 3D anchor tokens, forming a compact scaffold of the scene. Each token is then regressed into local Gaussians with learnable 3D offsets, decoupling primitive placement from input image grids. An Adaptive Token Expansion module predicts a token-level uncertainty score, supervised by rendering error maps, and selectively expands high-uncertainty tokens through learnable expansion layers. This sparse-to-adaptive formulation enables ATSplat to concentrate primitives in challenging regions while maintaining a compact representation. Experiments on two representative datasets, RealEstate10K and DL3DV, show that ATSplat achieves state-of-the-art rendering quality while reducing the number of Gaussians by more than 5.7times compared with dense feed-forward 3DGS methods. From 12 input images at 512 times 960 resolution, ATSplat completes reconstruction in less than a second using a single commercial GPU, and renders high-quality novel views at 1136 FPS (512 times 960) with only 311K Gaussians.