幾何学学習器としてのビデオ生成モデル
Video Generative Models as Geometry Learner
August 28, 2026
著者: Haosen Yang, Jifei Song, Zhensong Zhang, Xiatian Zhu, Jiankang Deng
cs.AI
要旨
最近の幾何推定への生成的アプローチは、事前学習済みの画像拡散モデルを適応させ、タスクを画像条件付き生成として扱う。既製の画像拡散モデルを活用するこれらの手法は、(i) 深度推定と表面法線推定のためのタスク固有の幾何モデルを個別に訓練し、これらの幾何ターゲットの本質的な相関を探る機会を失うか、あるいは (ii) 変更を加えた画像拡散バックボーン(例:改変された自己アテンション)を共同でファインチューニングするが、これは通常、大量のラベル付きデータを必要とする。これらの制限を原理的に克服するため、我々は事前学習済みのビデオ生成モデルを幾何推定のための統一されたデータ効率的なフレームワークとして再利用し、これを革新的に次フレーム予測タスクとして定式化する。我々の手法であるGeoNeXtは、ビデオモデルから自然に構造化された知識とより豊かな事前知識を継承しつつ、それらを画像と幾何ターゲットの共同モデリング(画像 ↔ 幾何)に適応させることで、よりデータ効率的かつ効果的な幾何学習を可能にする。広範な実験により、多様なデータセットにわたるゼロショット単眼深度推定と表面法線推定において我々の手法を検証し、大幅に少ないトレーニングデータでありながら、従来のタスク固有および統合生成の競合手法の両方を上回ることを示す。特筆すべきことに、我々の手法は100倍以上のデータで訓練された識別型の最先端手法に匹敵し、いくつかのベンチマークでは傑出した結果を示す。
English
Recent generative approaches to geometry estimation adapt pretrained image diffusion models and treat the task as image-conditioned generation. Leveraging off-the-shelf image diffusion models, they either (i) train task-specific geometry models (for depth and surface normal estimation) independently, losing the opportunity of exploring the intrinsic correlation of these geometric targets, or (ii) jointly fine-tune modified image diffusion backbones (e.g., altered self-attention), which typically demands substantial labeled data. To overcome these limitations in a principled fashion, we repurpose pretrained video generative models as a unified and data-efficient framework for geometry estimation, formulated innovatively as a next-frames prediction task. Our method, GeoNeXt, inherits naturally structured knowledge and richer priors from the video model, while further adapting them for joint modeling of images and geometry targets (image <-> geometry), enabling more data efficient and effective learning of geometry. Extensive experiments validate our method for zero-shot monocular depth and surface normal estimation across diverse datasets, outperforming both previous task-specific and unified generative competitors while using substantially less training data. Notably, our method rivals discriminative state-of-the-art approaches trained on over 100x more data and even standouts on several benchmarks.