ChatPaper.aiChatPaper

UniWorld-View: ビデオ拡散モデルによる大基線視点合成

UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Models

August 5, 2026
著者: Haiyang Zhou, Wangbo Yu, Chaoran Feng, Xunyu Zhou, Yonghong Tian, Li Yuan
cs.AI

要旨

ソーシャルメディア上で日常的に撮影された単眼ビデオや画像が豊富に存在することは、没入型コンテンツ制作における貴重な情報源であり、そのような疎な観測から新規視点を生成できれば、ユーザー体験を大幅に向上させることができる。しかし、入力の被覆が極めて限られている場合、精密なカメラ制御を伴う写実的かつ幾何的に一貫した視点生成は依然として困難である。NeRFや3D Gaussian Splatting (3DGS) などの再構成ベースの手法は、疎な入力条件下で性能が著しく低下し、遮蔽を明示的に扱うことができない。生成的手法はデータ要件を緩和する一方で、不正確または暗黙的な幾何学的ガイダンスに起因し、大基線長の視点合成には依然として課題が残る。これらの限界を克服するため、我々は単眼入力からの制御可能な大基線長新規視点合成のための統一フレームワークであるUniWorld-Viewを提案する。UniWorld-Viewは、明示的な3Dガイダンスと生成的拡散モデリングを統合し、精密なカメラ制御と幾何的に一貫した視点生成を実現する。幾何学的ガイダンスは、遮蔽を考慮した点群レンダリング戦略によって得られ、可視性の曖昧さを解消し、拡散ベースの合成のための正確な事前情報を提供する。このレンダリング戦略を強力なビデオ拡散バックボーンと組み合わせることで、UniWorld-Viewは極端なカメラ移動や広い基線長変化の下でも高忠実度の新規視点生成を達成し、さらに下流の動的3DGS再構成のためのマルチビュービデオを提供できる。WorldScoreベンチマークおよびゼロショットNVSベンチマークにおける実験により、UniWorld-Viewの制御可能性、幾何的一貫性、視覚的忠実度における有効性が示された。
English
The abundance of casually captured monocular videos and images on social media provides a valuable source for immersive content creation, where generating novel views from such sparse observations can greatly enhance user experiences. However, producing photorealistic and geometrically consistent views with precise camera control remains challenging when input coverage is extremely limited. Reconstruction-based approaches such as NeRF and 3D Gaussian Splatting (3DGS) deteriorate severely under sparse inputs and fail to explicitly handle occlusions. Generative methods ease data requirements but still struggle with large-baseline view synthesis due to inaccurate or implicit geometric guidance. To overcome these limitations, we introduce UniWorld-View, a unified framework for controllable large-baseline novel view synthesis from monocular inputs. UniWorld-View integrates explicit 3D guidance with generative diffusion modeling to enable precise camera control and geometrically consistent view generation. The geometric guidance is obtained through an occlusion-aware point cloud rendering strategy that resolves visibility ambiguities and provides accurate priors for diffusion-based synthesis. By coupling this rendering strategy with powerful video diffusion backbones, UniWorld-View achieves high-fidelity novel view generation even under extreme camera motions and wide-baseline changes, and can further provide multi-view videos for downstream dynamic 3DGS reconstruction. Experiments on the WorldScore benchmark and zero-shot NVS benchmarks demonstrate the effectiveness of UniWorld-View in controllability, geometric consistency, and visual fidelity.