ChatPaper.aiChatPaper

생성적 의미론적 장면 완성

Generative Semantic Scene Completion

August 27, 2026
저자: Shi Chen, Weifeng Ge
cs.AI

초록

야외 라이다 의미론적 장면 완성(SSC)은 대상 체적의 1%만 관측하는 스캔으로부터, 7,000배를 초과하는 클래스 불균형 하에서 밀집 의미론적 복셀 격자를 복원한다. 우리는 SSC를 생성적 의미론적 장면 완성(GSSC)으로 재구성한다: 세 가지 역할을 수행하는 단일 이산 확산 정식화. 첫째, 쌍을 이룬 희소-밀집 장면 합성(PS^3)은 정합된 희소 라이다 관측값과 그에 대응하는 밀집 의미론적 완성 결과를 생성하여, 긴 꼬리(long tail) 문제를 원천적으로 해결하고 학습 시 SemanticKITTI와 함께 사용하는 PS^3-SemanticKITTI 코퍼스를 산출한다. 둘째, 의미론 유도 생성적 장면 완성(SGSC)은 조감도 의미론적 지도와 희소 3D 특징 스트림을 통해 희소 스캔을 조건으로 삼아, 다항 이산 확산을 이용해 노이즈로부터 장면을 생성한다. 셋째, 동일한 프레임워크는 대신 구조화된 소스 이산 확산(S^2D^2)으로서 기존 완성 결과를 하나의 플로우 매칭 단계로 정제한다. S^2D^2는 기반 모델의 재학습이나 테스트 시 적응 없이 SGSC의 자체 출력과 테스트된 모든 외부 SSC 기반 모델의 mIoU를 향상시킨다. 가장 강력한 기반 모델에 대해 테스트 시 증강 없이 단일 단계만 적용하면 SemanticKITTI 비공개 테스트에서 38.8% mIoU를 달성한다. 우리가 아는 한, 이는 해당 리더보드에서 인과적(causal)이고 단일 스캔(single-sweep)이며 단일 샘플(single-sample) 조건을 만족하는 최고 결과로, 동일 제약 조건 하에서 이전에 발표된 최고 점수보다 +2.1퍼센트 포인트 높은 수치이다. 8개 시점 테스트 시 증강을 적용한 4회의 보정 단계는 해당 제약 조건 밖에서 39.2%에 도달한다.
English
Outdoor LiDAR semantic scene completion (SSC) recovers a dense semantic voxel grid from a scan observing 1% of the target volume, under class imbalance beyond 7,000x. We recast SSC as generative semantic scene completion (GSSC): a single discrete-diffusion formulation in three roles. First, paired sparse-dense scene synthesis (PS^3) generates matched sparse LiDAR observations with their dense semantic completions, addressing the long tail at its source and yielding the PS^3-SemanticKITTI corpus we train on alongside SemanticKITTI. Second, semantic-guided generative scene completion (SGSC) generates the scene from noise with multinomial discrete diffusion, conditioned on the sparse scan through a bird's-eye-view semantic map and a sparse 3D feature stream. Third, the same framework instead refines an existing completion in one flow-matching step: structured source discrete diffusion (S^2D^2). S^2D^2 improves the mIoU of SGSC's own output and every external SSC base tested, without base retraining or test-time adaptation. On the strongest base, one step without test-time augmentation reaches 38.8% mIoU on the SemanticKITTI hidden test. To our knowledge that is the best causal, single-sweep, single-sample result on that leaderboard, +2.1 pp over the previous best published score under the same restriction. Four correction steps with eight-view test-time augmentation reach 39.2%, outside that restriction.