ChatPaper.aiChatPaper

CGGS: 一貫性拡張幾何学的ガウシアンスプラッティングによるエゴセントリック3Dシーン生成

CGGS: Consistency-Augmented Geometric Gaussian Splatting for Ego-centric 3D Scene Generation

July 4, 2026
著者: Zhenyu Sun, Xiaohan Zhang, Qi Liu, Huan Wang
cs.AI

要旨

自己中心的な3Dシーン生成においては、視点の重なりが限定的であることや、シーン解釈に対する個人の視点の支配的な影響により、課題が残されている。これらの要因は、視点一貫性があり意味的に整合した視覚コンテンツの生成や、正確な幾何学的構造の構築を妨げる。本稿では、3Dコンテンツ認識を向上させ、自己中心的なシーン生成における幾何学的歪みに対処することを目的としたテキストから3Dへのフレームワーク「CGGS」を提案する。まず、自己中心生成器(Ego-centric Generator)を提案する。これは、一貫性強化損失を用いてマルチビュー潜在拡散モデルをファインチューニングし、テキスト記述に整合した一貫性のある高忠実度の2Dコンテンツを生成する。次に、レイアウトデコレータ(Layout Decorator)は、オプティカルフローとポイントトラック対応を活用して深度を推定し、自己中心的な2D事前情報から粗いレイアウトとしての密な点群を生成する。この初期化に基づき、幾何学的リファイナ(Geometric Refiner)を提案する。これは、エントロピーベースの相互情報量深度損失(MID)と、視覚品質と幾何学的構造を改善するための階層的最適化スキームを組み合わせて、3Dガウシアン再構成を強化する。包括的な実験により、CGGSはテキスト駆動型の一貫性のある正確な3Dシーンの生成において、従来手法を上回ることが示された。プロジェクトページ: https://cggs-26.github.io/cggs26/
English
Challenges remain in ego-centric 3D scene generation due to limited view overlap and the dominant influence of individual perspectives on scene interpretation. These factors hinder the creation of viewpoint-consistent and semantically aligned visual content, as well as the construction of accurate geometric structures. In this paper, we propose CGGS, a text-to-3D framework aiming to enhance 3D-content-awareness and address geometric distortions in ego-centric scene generation. Firstly, the Ego-centric Generator is proposed by fine-tuning a Multi-View Latent Diffusion Model with consistency-augmented loss to generate consistent, high-fidelity 2D content aligned with textual descriptions. Then, Layout Decorator leverages optical flow and point-track correspondence to estimate depth, therefore producing dense point clouds as coarse layouts from the ego-centric 2D priors. Building on this initialization, Geometric Refiner is proposed to enhance 3D Gaussian reconstruction via an entropy-based Mutual Information Depth Loss (MID) combined with a hierarchical optimization scheme for improving visual quality and geometric structure. Comprehensive experiments demonstrate that softred{CGGS} outperforms previous methods in generating coherent and accurate text-driven 3D scenes. Project page: https://cggs-26.github.io/cggs26/.