視訊生成模型作為幾何學習器
Video Generative Models as Geometry Learner
August 28, 2026
作者: Haosen Yang, Jifei Song, Zhensong Zhang, Xiatian Zhu, Jiankang Deng
cs.AI
摘要
近期針對幾何估計的生成式方法會調整預訓練的影像擴散模型,並將任務視為影像條件生成。透過利用現成的影像擴散模型,這類方法要麼 (i) 獨立訓練特定任務的幾何模型(用於深度與表面法向量估計),因而失去探索這些幾何目標內在關聯的機會;要麼 (ii) 聯合微調修改後的影像擴散主幹網路(例如改變自注意力機制),這通常需要大量標註資料。為了以有原則的方式克服這些限制,我們將預訓練的影片生成模型重新定位為一個統一且資料高效的幾何估計框架,並創新地將此任務表述為下一幀預測。我們的方法 GeoNeXt 繼承了影片模型中自然結構化的知識與更豐富的先驗,同時進一步調整它們以聯合建模影像與幾何目標(影像 <-> 幾何),從而實現更具資料效率且更有效的幾何學習。大量實驗驗證了我們的方法在各種資料集上進行零樣本單目深度與表面法向量估計的效能,不僅優於先前的任務特定與統一生成式競爭方法,且使用的訓練資料大幅減少。值得注意的是,我們的方法能夠與使用超過100倍資料訓練的判別式最先進方法相抗衡,甚至在數個基準上脫穎而出。
English
Recent generative approaches to geometry estimation adapt pretrained image diffusion models and treat the task as image-conditioned generation. Leveraging off-the-shelf image diffusion models, they either (i) train task-specific geometry models (for depth and surface normal estimation) independently, losing the opportunity of exploring the intrinsic correlation of these geometric targets, or (ii) jointly fine-tune modified image diffusion backbones (e.g., altered self-attention), which typically demands substantial labeled data. To overcome these limitations in a principled fashion, we repurpose pretrained video generative models as a unified and data-efficient framework for geometry estimation, formulated innovatively as a next-frames prediction task. Our method, GeoNeXt, inherits naturally structured knowledge and richer priors from the video model, while further adapting them for joint modeling of images and geometry targets (image <-> geometry), enabling more data efficient and effective learning of geometry. Extensive experiments validate our method for zero-shot monocular depth and surface normal estimation across diverse datasets, outperforming both previous task-specific and unified generative competitors while using substantially less training data. Notably, our method rivals discriminative state-of-the-art approaches trained on over 100x more data and even standouts on several benchmarks.