ChatPaper.aiChatPaper

UniWorld-View:基於視頻擴散模型的大基線視圖合成

UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Models

August 5, 2026
作者: Haiyang Zhou, Wangbo Yu, Chaoran Feng, Xunyu Zhou, Yonghong Tian, Li Yuan
cs.AI

摘要

社群媒體上隨手捕捉的單目影片與影像數量豐富,為沉浸式內容創作提供了有價值的來源,其中從此類稀疏觀測生成新視圖,能大幅提升使用者體驗。然而,當輸入覆蓋範圍極度有限時,要產生具有照片寫實度與幾何一致性的視圖,並具備精確的相機控制,仍具挑戰性。基於重建的方法(如 NeRF 與 3D 高斯潑濺(3DGS))在稀疏輸入下會嚴重退化,且無法明確處理遮擋。生成式方法雖降低了資料需求,但因幾何引導不準確或具隱式性,在大基線視圖合成上仍顯吃力。為克服這些限制,我們提出 UniWorld-View,一個用於從單目輸入進行可控大基線新視圖合成的統一框架。UniWorld-View 將顯式 3D 幾何引導與生成式擴散建模整合,以實現精確的相機控制與幾何一致的視圖生成。幾何引導是透過一種具遮擋感知的點雲渲染策略取得,該策略解決了可見性歧義,並為基於擴散的合成提供準確的先驗。透過將此渲染策略與強大的視訊擴散骨幹網路結合,UniWorld-View 即使在極端的相機運動與寬基線變化下,也能生成高保真的新視圖,並可進一步提供多視圖影片,以供下游的動態 3DGS 重建使用。在 WorldScore 基準與零樣本新視圖合成基準上的實驗,證明了 UniWorld-View 在可控性、幾何一致性與視覺保真度方面的有效性。
English
The abundance of casually captured monocular videos and images on social media provides a valuable source for immersive content creation, where generating novel views from such sparse observations can greatly enhance user experiences. However, producing photorealistic and geometrically consistent views with precise camera control remains challenging when input coverage is extremely limited. Reconstruction-based approaches such as NeRF and 3D Gaussian Splatting (3DGS) deteriorate severely under sparse inputs and fail to explicitly handle occlusions. Generative methods ease data requirements but still struggle with large-baseline view synthesis due to inaccurate or implicit geometric guidance. To overcome these limitations, we introduce UniWorld-View, a unified framework for controllable large-baseline novel view synthesis from monocular inputs. UniWorld-View integrates explicit 3D guidance with generative diffusion modeling to enable precise camera control and geometrically consistent view generation. The geometric guidance is obtained through an occlusion-aware point cloud rendering strategy that resolves visibility ambiguities and provides accurate priors for diffusion-based synthesis. By coupling this rendering strategy with powerful video diffusion backbones, UniWorld-View achieves high-fidelity novel view generation even under extreme camera motions and wide-baseline changes, and can further provide multi-view videos for downstream dynamic 3DGS reconstruction. Experiments on the WorldScore benchmark and zero-shot NVS benchmarks demonstrate the effectiveness of UniWorld-View in controllability, geometric consistency, and visual fidelity.