저중첩 캡처로부터의 4D 인간-장면 재구성

4D Human-Scene Reconstruction from Low-Overlap Captures

July 10, 2026
저자: Minhyuk Hwang, Sangmin Kim, Seunguk Do, Daneul Kim, Jaesik Park
cs.AI

초록

기존의 동적 인간 퍼포먼스 체적 캡처 기술은 밀집된 카메라 배열을 통해 높은 정확도를 달성합니다. 그러나 실제 환경에서는 중첩이 낮은 소수의 카메라만 사용 가능하여 출력 품질이 저하되고 넓은 영역이 관측되지 않은 상태로 남습니다. 최근의 4D 재구성 방법은 저중첩 환경에 초점을 맞추고 있지만, 관측이 부족한 영역에서 여전히 눈에 띄는 인공물이 발생합니다. 비디오 확산 모델이 또 다른 대안으로 부상했으나, 인간에 대해서는 기하학적으로 일관성 없는 결과를 보여줍니다. 이러한 한계를 극복하기 위해, 우리는 배경과 인간을 분리하여 희소하고 중첩이 낮은 카메라로부터 4D 인간 장면을 재구성하는 파이프라인인 StudioRecon을 제안합니다. 비디오 확산 모델을 사용하여 수백 개의 카메라 제어 가능한 새로운 시점을 합성함으로써 배경 감독을 밀집화합니다. 또한, 시점 간 신원 연관 및 삼각 측정 기반 다중 시점 키포인트 피팅을 통해 변형 가능한 가우시안 휴먼을 강건하게 초기화합니다. 마지막으로, 움직임 적응 일관성 주입을 갖춘 재귀적 향상 모듈이 구성된 출력을 조화시켜 남아 있는 인공물을 추가로 방지합니다. 우리는 네 개의 실제 데이터셋에서 최첨단 새로운 시점 합성 성능을 달성했으며, 새로운 궤적 렌더링 및 인간 대체와 같은 응용을 시연합니다.
English
Existing volumetric capture of dynamic human performance achieves high fidelity with dense camera arrays. However, in real-world scenarios, only a handful of low-overlap cameras are available, which degrades the output quality and leaves large areas unobserved. Recent 4D reconstruction methods have focused on low-overlap settings, yet they still produce noticeable artifacts in under-observed regions. Video diffusion models have emerged as another option, but they show geometrically inconsistent results for humans. To address these limitations, we propose StudioRecon, a pipeline that reconstructs 4D human scenes from sparse, low-overlap cameras by decoupling background and humans. We densify background supervision by synthesizing hundreds of camera-controlled novel views with a video diffusion model. We also robustly initialize deformable Gaussian humans with cross-view identity association and triangulated multi-view keypoint fitting. Finally, our recursive enhancement module with motion-adaptive consistency injection harmonizes the composed output, thereby further avoiding remaining artifacts. We achieve state-of-the-art novel view synthesis across four real-world datasets and demonstrate applications such as novel trajectory rendering and human replacement.
PDF412July 15, 2026