ChatPaper.aiChatPaper

生成的セマンティックシーン補完

Generative Semantic Scene Completion

August 27, 2026
著者: Shi Chen, Weifeng Ge
cs.AI

要旨

屋外LiDARセマンティックシーン補完(SSC)は、対象ボリュームの1%しか観測しないスキャンから、7,000倍を超えるクラス不均衡のもとで、高密度なセマンティックボクセルグリッドを復元する。我々はSSCを生成セマンティックシーン補完(GSSC)として再定義する。これは単一の離散拡散定式化であり、3つの役割を担う。まず、ペア型スパース・デンスシーン合成(PS^3)は、対応するスパースLiDAR観測とその高密度セマンティック補完を生成し、ロングテール問題をその発生源で対処し、SemanticKITTIと併用して学習するPS^3-SemanticKITTIコーパスを提供する。次に、セマンティック誘導生成シーン補完(SGSC)は、鳥瞰図セマンティックマップとスパース3D特徴ストリームを介してスパーススキャンに条件付けられ、多項離散拡散によりノイズからシーンを生成する。第3に、同じフレームワークは、1回のフローマッチングステップで既存の補完結果を改良する。これが構造化ソース離散拡散(S^2D^2)である。S^2D^2は、ベースの再学習やテスト時適応を必要とせずに、SGSC自身の出力と、テストしたすべての外部SSCベースのmIoUを向上させる。最強のベースでは、テスト時拡張なしの1ステップで、SemanticKITTI非公開テストにおいて38.8%のmIoUに達する。我々の知る限り、これはそのリーダーボードにおける因果的・単一スイープ・単一サンプルでの最良の結果であり、同じ制約下での従来の公開最高スコアを+2.1ポイント上回る。8視点のテスト時拡張を用いた4回の補正ステップでは、その制約の外で39.2%に達する。
English
Outdoor LiDAR semantic scene completion (SSC) recovers a dense semantic voxel grid from a scan observing 1% of the target volume, under class imbalance beyond 7,000x. We recast SSC as generative semantic scene completion (GSSC): a single discrete-diffusion formulation in three roles. First, paired sparse-dense scene synthesis (PS^3) generates matched sparse LiDAR observations with their dense semantic completions, addressing the long tail at its source and yielding the PS^3-SemanticKITTI corpus we train on alongside SemanticKITTI. Second, semantic-guided generative scene completion (SGSC) generates the scene from noise with multinomial discrete diffusion, conditioned on the sparse scan through a bird's-eye-view semantic map and a sparse 3D feature stream. Third, the same framework instead refines an existing completion in one flow-matching step: structured source discrete diffusion (S^2D^2). S^2D^2 improves the mIoU of SGSC's own output and every external SSC base tested, without base retraining or test-time adaptation. On the strongest base, one step without test-time augmentation reaches 38.8% mIoU on the SemanticKITTI hidden test. To our knowledge that is the best causal, single-sweep, single-sample result on that leaderboard, +2.1 pp over the previous best published score under the same restriction. Four correction steps with eight-view test-time augmentation reach 39.2%, outside that restriction.