動画生成モデルは汎用視覚学習器である

Video Generation Models are General-Purpose Vision Learners

July 10, 2026
著者: Letian Wang, Chuhan Zhang, Rishabh Kabra, Jasper Uijlings, Steven Waslander, Andrew Zisserman, Joao Carreira, Kaiming He, Misha Andriluka, Eduard Gabriel Bazavan, Andrei Zanfir, Cristian Sminchisescu
cs.AI

要旨

次トークン予測によって駆動され、NLPはタスク固有モデルから強力な汎用基盤モデルへと移行した。では、コンピュータビジョンにおいて汎用モデルを実現するための同等の触媒とは何だろうか。本論文では、大規模テキストから動画への生成がコンピュータビジョンにおける強力な事前学習パラダイムとして機能し、汎用的な視覚知能に必要な時空間事前知識、視覚-言語アライメント、およびスケーラビリティを提供すると主張する。我々はGenCeptionを提案する。これは、事前学習された動画生成拡散バックボーンを活用し、テキスト指示によって様々な視覚タスクを実行可能なフィードフォワード認識モデルを定義する。実験結果は、GenCeptionが奥行き、法線、カメラポーズ推定、表現参照セグメンテーション、3Dキーポイント予測など多様なタスクにおいて最先端の性能を達成し、しばしば専門モデル(例:DepthAnything3、SAM3、D4RT、VGGT-Omega、Sapiens、David、Genmo、Lotus-2)に匹敵または凌駕することを示す。さらに、動画生成事前学習バックボーンは、同等設定下で代替事前学習パラダイム(例:V-JEPA、Video MAE)を上回る。重要なことに、GenCeptionはデータおよびモデルのスケーリング特性とともに、優れたデータ効率を示し、D4RTやVGGT-Omegaなどの主要モデルと比較して7~500分の1の訓練データで同等の性能を達成する。最後に、GenCeptionは興味深い創発的行動も示す。合成人間動画のみで訓練されたモデルが実世界の映像や分布外の物体カテゴリ(例:動物やロボット)に一般化する。これらの発見は、動画生成が単なる合成ツールではなく、物理世界における汎用視覚知能への基盤的経路であることを示唆する。プロジェクトページ: https://genception.github.io
English
Driven by next-token prediction, NLP shifted from task-specific models into powerful generalist foundation models. What, then, is the equivalent catalyst needed to achieve a general-purpose model in computer vision? In this paper, we contend that large-scale text-to-video generation serves as a strong pre-training paradigm for computer vision, providing the necessary spatiotemporal priors, vision-language alignment, and scalability required for general visual intelligence. We introduce GenCeption, which leverages a pre-trained video generative diffusion backbone to define a feed-forward perception model, capable of performing various vision tasks steered by text instructions. Empirical results demonstrate that GenCeption achieves state-of-the-art performance across a diverse suite of tasks, including depth, surface normal, and camera pose estimation, expression-referring segmentation, and 3D keypoint prediction, often matching or surpassing specialized models (e.g. DepthAnything3, SAM3, D4RT, VGGT-Omega, Sapiens, David, Genmo, and Lotus-2). Furthermore, the video generative pretrained backbone outperforms alternative pretraining paradigms (e.g., V-JEPA, and Video MAE) under comparable settings. Importantly, GenCeption exhibits preliminary data and model scaling properties along with exceptional data efficiency, where it achieves comparable performance with leading models like D4RT and VGGT-Omega with 7 to 500 less training data. Finally, GenCeption also exhibits intriguing emergent behaviors: a model trained exclusively on synthetic human videos generalizes to real-world footage and out-of-distribution object categories (e.g., animals and robots). These findings suggest that video generation is not merely a synthesis tool, but a foundational path toward generalist vision intelligence for the physical world. Project page: https://genception.github.io
PDF400July 14, 2026