低重叠度捕获下的4D人体-场景重建
4D Human-Scene Reconstruction from Low-Overlap Captures
July 10, 2026
作者: Minhyuk Hwang, Sangmin Kim, Seunguk Do, Daneul Kim, Jaesik Park
cs.AI
摘要
现有的动态人体体积捕捉技术虽然能通过密集相机阵列实现高保真度,但在现实场景中通常仅有少量低重叠相机可用,这会导致输出质量下降并留下大量未观测区域。近年来,4D重建方法虽聚焦于低重叠场景,但在未充分观测区域仍会产生明显伪影。视频扩散模型虽成为另一种选择,但应用于人体时会出现几何不一致的结果。为解决这些限制,我们提出StudioRecon——一种通过解耦背景与人体,从稀疏低重叠相机重建4D人体场景的流水线。我们利用视频扩散模型合成数百个受相机参数控制的新视角,从而增强背景监督信号。同时,我们结合跨视角身份关联与三角化多视角关键点拟合,稳健初始化可变形高斯人体模型。最后,我们提出的递归增强模块通过运动自适应一致性注入对合成输出进行调和,进一步消除残留伪影。在四个真实场景数据集上,我们实现了当前最优的新视角合成效果,并展示了新颖轨迹渲染与人体替换等应用。
English
Existing volumetric capture of dynamic human performance achieves high fidelity with dense camera arrays. However, in real-world scenarios, only a handful of low-overlap cameras are available, which degrades the output quality and leaves large areas unobserved. Recent 4D reconstruction methods have focused on low-overlap settings, yet they still produce noticeable artifacts in under-observed regions. Video diffusion models have emerged as another option, but they show geometrically inconsistent results for humans. To address these limitations, we propose StudioRecon, a pipeline that reconstructs 4D human scenes from sparse, low-overlap cameras by decoupling background and humans. We densify background supervision by synthesizing hundreds of camera-controlled novel views with a video diffusion model. We also robustly initialize deformable Gaussian humans with cross-view identity association and triangulated multi-view keypoint fitting. Finally, our recursive enhancement module with motion-adaptive consistency injection harmonizes the composed output, thereby further avoiding remaining artifacts. We achieve state-of-the-art novel view synthesis across four real-world datasets and demonstrate applications such as novel trajectory rendering and human replacement.