ChatPaper.aiChatPaper

UniWorld-View:基于视频扩散模型的大基线视角合成

UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Models

August 5, 2026
作者: Haiyang Zhou, Wangbo Yu, Chaoran Feng, Xunyu Zhou, Yonghong Tian, Li Yuan
cs.AI

摘要

社交媒体上随手拍摄的单目视频和图像数量庞大,为沉浸式内容创作提供了宝贵来源,从这类稀疏观测中生成新视角可以极大提升用户体验。然而,当输入覆盖范围极为有限时,生成具有精确相机控制的照片级真实感且几何一致的视角仍然具有挑战性。基于重建的方法(如NeRF和3D高斯泼溅(3DGS))在稀疏输入下会严重退化,并且无法显式处理遮挡。生成方法放宽了数据需求,但由于几何引导不准确或隐式,仍然难以应对大基线视角合成。为了克服这些限制,我们提出了UniWorld-View,一个统一的框架,用于从单目输入进行可控的大基线新视角合成。UniWorld-View将显式3D引导与生成式扩散建模相结合,以实现精确的相机控制和几何一致的新视角生成。该几何引导通过一种遮挡感知的点云渲染策略获得,该策略解决了可见性歧义,并为基于扩散的合成提供了准确的先验。通过将该渲染策略与强大的视频扩散主干相结合,UniWorld-View即使在极端相机运动和大基线变化下也能实现高保真的新视角生成,并进一步为下游动态3DGS重建提供多视角视频。在WorldScore基准和零样本NVS基准上的实验证明了UniWorld-View在可控性、几何一致性和视觉保真度方面的有效性。
English
The abundance of casually captured monocular videos and images on social media provides a valuable source for immersive content creation, where generating novel views from such sparse observations can greatly enhance user experiences. However, producing photorealistic and geometrically consistent views with precise camera control remains challenging when input coverage is extremely limited. Reconstruction-based approaches such as NeRF and 3D Gaussian Splatting (3DGS) deteriorate severely under sparse inputs and fail to explicitly handle occlusions. Generative methods ease data requirements but still struggle with large-baseline view synthesis due to inaccurate or implicit geometric guidance. To overcome these limitations, we introduce UniWorld-View, a unified framework for controllable large-baseline novel view synthesis from monocular inputs. UniWorld-View integrates explicit 3D guidance with generative diffusion modeling to enable precise camera control and geometrically consistent view generation. The geometric guidance is obtained through an occlusion-aware point cloud rendering strategy that resolves visibility ambiguities and provides accurate priors for diffusion-based synthesis. By coupling this rendering strategy with powerful video diffusion backbones, UniWorld-View achieves high-fidelity novel view generation even under extreme camera motions and wide-baseline changes, and can further provide multi-view videos for downstream dynamic 3DGS reconstruction. Experiments on the WorldScore benchmark and zero-shot NVS benchmarks demonstrate the effectiveness of UniWorld-View in controllability, geometric consistency, and visual fidelity.