CGGS: 자아 중심 3D 장면 생성을 위한 일관성 강화 기하학적 가우시안 스플래팅
CGGS: Consistency-Augmented Geometric Gaussian Splatting for Ego-centric 3D Scene Generation
July 4, 2026
저자: Zhenyu Sun, Xiaohan Zhang, Qi Liu, Huan Wang
cs.AI
초록
자기중심적(ego-centric) 3D 장면 생성에는 제한된 시점 중첩과 개별 시점이 장면 해석에 미치는 지배적 영향으로 인해 어려움이 남아 있다. 이러한 요인들은 시점 일관성 및 의미론적 정렬을 갖춘 시각적 콘텐츠의 생성을 저해할 뿐만 아니라 정확한 기하학적 구조의 구축을 방해한다. 본 논문에서는 3D 콘텐츠 인지도를 향상시키고 자기중심적 장면 생성에서의 기하학적 왜곡을 해결하기 위해 텍스트-3D 프레임워크인 CGGS를 제안한다. 첫째, 자기중심적 생성기(Ego-centric Generator)는 일관성 증강 손실로 다중 뷰 잠재 확산 모델을 미세 조정하여 텍스트 설명과 정렬된 일관되고 고충실도의 2D 콘텐츠를 생성하도록 설계된다. 다음으로, 레이아웃 데코레이터(Layout Decorator)는 광학 흐름과 점 추적 대응을 활용하여 깊이를 추정함으로써 자기중심적 2D 사전 정보로부터 대략적인 레이아웃으로서의 밀집 점군을 생성한다. 이러한 초기화를 기반으로, 기하학적 정제기(Geometric Refiner)는 엔트로피 기반 상호 정보 깊이 손실(MID)과 시각적 품질 및 기하학적 구조를 개선하기 위한 계층적 최적화 기법을 결합하여 3D 가우시안 재구성을 향상시킨다. 포괄적인 실험을 통해 CGGS가 텍스트 기반의 일관되고 정확한 3D 장면 생성에서 이전 방법들을 능가함을 입증한다. 프로젝트 페이지: https://cggs-26.github.io/cggs26/.
English
Challenges remain in ego-centric 3D scene generation due to limited view overlap and the dominant influence of individual perspectives on scene interpretation. These factors hinder the creation of viewpoint-consistent and semantically aligned visual content, as well as the construction of accurate geometric structures. In this paper, we propose CGGS, a text-to-3D framework aiming to enhance 3D-content-awareness and address geometric distortions in ego-centric scene generation. Firstly, the Ego-centric Generator is proposed by fine-tuning a Multi-View Latent Diffusion Model with consistency-augmented loss to generate consistent, high-fidelity 2D content aligned with textual descriptions. Then, Layout Decorator leverages optical flow and point-track correspondence to estimate depth, therefore producing dense point clouds as coarse layouts from the ego-centric 2D priors. Building on this initialization, Geometric Refiner is proposed to enhance 3D Gaussian reconstruction via an entropy-based Mutual Information Depth Loss (MID) combined with a hierarchical optimization scheme for improving visual quality and geometric structure. Comprehensive experiments demonstrate that softred{CGGS} outperforms previous methods in generating coherent and accurate text-driven 3D scenes. Project page: https://cggs-26.github.io/cggs26/.