低重疊捕捉下的四維人與場景重建
4D Human-Scene Reconstruction from Low-Overlap Captures
July 10, 2026
作者: Minhyuk Hwang, Sangmin Kim, Seunguk Do, Daneul Kim, Jaesik Park
cs.AI
摘要
現有的動態人體體積捕捉技術透過密集的相機陣列可達到高保真度,但在實際場景中,僅有少數低重疊相機可供使用,這會降低輸出品質並留下大片未觀測區域。近期的4D重建方法雖已針對低重疊設定進行研究,但在觀測不足的區域仍會產生明顯的偽影。影片擴散模型雖然成為另一種選擇,但其在幾何一致性上對人體表現不佳。為解決這些限制,我們提出StudioRecon,這是一種透過分離背景與人體,從稀疏低重疊相機重建4D人體場景的流程。我們透過影片擴散模型合成數百個受相機控制的新視角,以稠密化背景監督。同時,我們利用跨視角身分關聯與三角化多視角關鍵點擬合,穩健地初始化可形變高斯人體。最後,我們提出的遞迴增強模組搭配運動自適應一致性注入,協調合成輸出,進一步避免殘留偽影。我們在四個真實世界資料集上達成最先進的新視角合成,並展示新軌跡渲染與人體替換等應用。
English
Existing volumetric capture of dynamic human performance achieves high fidelity with dense camera arrays. However, in real-world scenarios, only a handful of low-overlap cameras are available, which degrades the output quality and leaves large areas unobserved. Recent 4D reconstruction methods have focused on low-overlap settings, yet they still produce noticeable artifacts in under-observed regions. Video diffusion models have emerged as another option, but they show geometrically inconsistent results for humans. To address these limitations, we propose StudioRecon, a pipeline that reconstructs 4D human scenes from sparse, low-overlap cameras by decoupling background and humans. We densify background supervision by synthesizing hundreds of camera-controlled novel views with a video diffusion model. We also robustly initialize deformable Gaussian humans with cross-view identity association and triangulated multi-view keypoint fitting. Finally, our recursive enhancement module with motion-adaptive consistency injection harmonizes the composed output, thereby further avoiding remaining artifacts. We achieve state-of-the-art novel view synthesis across four real-world datasets and demonstrate applications such as novel trajectory rendering and human replacement.