视频生成模型作为几何学习器
Video Generative Models as Geometry Learner
August 28, 2026
作者: Haosen Yang, Jifei Song, Zhensong Zhang, Xiatian Zhu, Jiankang Deng
cs.AI
摘要
最近的几何估计生成方法利用预训练图像扩散模型,将任务视为图像条件生成。利用现成的图像扩散模型,它们要么(i)独立训练任务特定的几何模型(用于深度和表面法线估计),从而失去了探索这些几何目标内在关联的机会;要么(ii)联合微调修改过的图像扩散主干网络(例如改变自注意力),但这通常需要大量标注数据。为了从原理上克服这些限制,我们将预训练视频生成模型重新利用为统一且数据高效的几何估计框架,创新性地将其表述为下一帧预测任务。我们的方法 GeoNeXt 继承了视频模型中天然的结构化知识和更丰富的先验,同时进一步调整它们以适应图像与几何目标(图像 <-> 几何)的联合建模,使几何学习更加数据高效且有效。大量实验验证了我们的方法在多样数据集上的零样本单目深度和表面法线估计能力,在使用显著更少训练数据的情况下,超越了以往任务特定和统一生成式的对比方法。值得注意的是,我们的方法可与在超过100倍数据上训练的判别式最先进方法相媲美,甚至在多个基准上表现突出。
English
Recent generative approaches to geometry estimation adapt pretrained image diffusion models and treat the task as image-conditioned generation. Leveraging off-the-shelf image diffusion models, they either (i) train task-specific geometry models (for depth and surface normal estimation) independently, losing the opportunity of exploring the intrinsic correlation of these geometric targets, or (ii) jointly fine-tune modified image diffusion backbones (e.g., altered self-attention), which typically demands substantial labeled data. To overcome these limitations in a principled fashion, we repurpose pretrained video generative models as a unified and data-efficient framework for geometry estimation, formulated innovatively as a next-frames prediction task. Our method, GeoNeXt, inherits naturally structured knowledge and richer priors from the video model, while further adapting them for joint modeling of images and geometry targets (image <-> geometry), enabling more data efficient and effective learning of geometry. Extensive experiments validate our method for zero-shot monocular depth and surface normal estimation across diverse datasets, outperforming both previous task-specific and unified generative competitors while using substantially less training data. Notably, our method rivals discriminative state-of-the-art approaches trained on over 100x more data and even standouts on several benchmarks.