CGGS:用於自我中心3D場景生成的一致性增強幾何高斯潑濺
CGGS: Consistency-Augmented Geometric Gaussian Splatting for Ego-centric 3D Scene Generation
July 4, 2026
作者: Zhenyu Sun, Xiaohan Zhang, Qi Liu, Huan Wang
cs.AI
摘要
以自我为中心的3D场景生成仍面临挑战,原因在于视图重叠有限,且个体视角对场景解读具有主导性影响。这些因素阻碍了生成视角一致且语义对齐的视觉内容,以及构建准确的几何结构。本文提出CGGS——一个文本到3D的框架,旨在提升3D内容感知能力并解决以自我为中心的场景生成中的几何畸变问题。首先,我们提出以自我为中心的生成器,通过采用一致性增强损失对多视图潜在扩散模型进行微调,以生成与文本描述对齐的一致、高保真2D内容。然后,布局装饰器利用光流和点轨迹对应关系估计深度,从而从以自我为中心的2D先验中生成密集点云作为粗略布局。在此初始化的基础上,我们提出几何精炼器,通过基于熵的互信息深度损失(MID)结合分层优化方案,提升3D高斯重建的视觉质量与几何结构。大量实验表明,CGGS在生成连贯且准确的文本驱动3D场景方面优于先前方法。项目页面:https://cggs-26.github.io/cggs26/。
English
Challenges remain in ego-centric 3D scene generation due to limited view overlap and the dominant influence of individual perspectives on scene interpretation. These factors hinder the creation of viewpoint-consistent and semantically aligned visual content, as well as the construction of accurate geometric structures. In this paper, we propose CGGS, a text-to-3D framework aiming to enhance 3D-content-awareness and address geometric distortions in ego-centric scene generation. Firstly, the Ego-centric Generator is proposed by fine-tuning a Multi-View Latent Diffusion Model with consistency-augmented loss to generate consistent, high-fidelity 2D content aligned with textual descriptions. Then, Layout Decorator leverages optical flow and point-track correspondence to estimate depth, therefore producing dense point clouds as coarse layouts from the ego-centric 2D priors. Building on this initialization, Geometric Refiner is proposed to enhance 3D Gaussian reconstruction via an entropy-based Mutual Information Depth Loss (MID) combined with a hierarchical optimization scheme for improving visual quality and geometric structure. Comprehensive experiments demonstrate that softred{CGGS} outperforms previous methods in generating coherent and accurate text-driven 3D scenes. Project page: https://cggs-26.github.io/cggs26/.