UniWorld-View: 비디오 확산 모델을 통한 대규모 베이스라인 시점 합성
UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Models
August 5, 2026
저자: Haiyang Zhou, Wangbo Yu, Chaoran Feng, Xunyu Zhou, Yonghong Tian, Li Yuan
cs.AI
초록
소셜 미디어에는 일상적으로 촬영된 단안 비디오와 이미지가 풍부하며, 이는 몰입형 콘텐츠 제작에 귀중한 원천을 제공한다. 이러한 희소한 관측으로부터 새로운 시점을 생성하는 것은 사용자 경험을 크게 향상시킬 수 있다. 그러나 입력 커버리지가 극도로 제한된 경우에는 정밀한 카메라 제어와 함께 사진처럼 사실적이고 기하학적으로 일관된 시점을 생성하는 것이 여전히 어려운 과제로 남아 있다. NeRF 및 3D 가우시안 스플래팅(3DGS)과 같은 재구성 기반 접근법은 희소 입력 조건에서 성능이 심각하게 저하되며 폐색을 명시적으로 처리하지 못한다. 생성 기반 방법은 데이터 요구 사항을 완화하지만, 부정확하거나 암시적인 기하학적 가이던스로 인해 큰 베이스라인 시점 합성에는 여전히 어려움을 겪는다. 이러한 한계를 극복하기 위해, 우리는 단안 입력으로부터 제어 가능한 큰 베이스라인 새로운 시점 합성을 위한 통합 프레임워크인 UniWorld-View를 제안한다. UniWorld-View는 명시적 3D 가이던스와 생성 확산 모델링을 통합하여 정밀한 카메라 제어와 기하학적으로 일관된 시점 생성을 가능하게 한다. 기하학적 가이던스는 가시성 모호성을 해결하고 확산 기반 합성을 위한 정확한 사전 정보를 제공하는 폐색 인식 포인트 클라우드 렌더링 전략을 통해 얻어진다. 이 렌더링 전략을 강력한 비디오 확산 백본과 결합함으로써, UniWorld-View는 극단적인 카메라 움직임과 넓은 베이스라인 변화에서도 고충실도 새로운 시점 생성을 달성하며, 다운스트림 동적 3DGS 재구성을 위한 다시점 비디오도 추가로 제공할 수 있다. WorldScore 벤치마크와 제로샷 NVS 벤치마크에서의 실험은 제어 가능성, 기하학적 일관성, 시각적 충실도 측면에서 UniWorld-View의 효과성을 입증한다.
English
The abundance of casually captured monocular videos and images on social media provides a valuable source for immersive content creation, where generating novel views from such sparse observations can greatly enhance user experiences. However, producing photorealistic and geometrically consistent views with precise camera control remains challenging when input coverage is extremely limited. Reconstruction-based approaches such as NeRF and 3D Gaussian Splatting (3DGS) deteriorate severely under sparse inputs and fail to explicitly handle occlusions. Generative methods ease data requirements but still struggle with large-baseline view synthesis due to inaccurate or implicit geometric guidance. To overcome these limitations, we introduce UniWorld-View, a unified framework for controllable large-baseline novel view synthesis from monocular inputs. UniWorld-View integrates explicit 3D guidance with generative diffusion modeling to enable precise camera control and geometrically consistent view generation. The geometric guidance is obtained through an occlusion-aware point cloud rendering strategy that resolves visibility ambiguities and provides accurate priors for diffusion-based synthesis. By coupling this rendering strategy with powerful video diffusion backbones, UniWorld-View achieves high-fidelity novel view generation even under extreme camera motions and wide-baseline changes, and can further provide multi-view videos for downstream dynamic 3DGS reconstruction. Experiments on the WorldScore benchmark and zero-shot NVS benchmarks demonstrate the effectiveness of UniWorld-View in controllability, geometric consistency, and visual fidelity.