影片生成模型是通用視覺學習器
Video Generation Models are General-Purpose Vision Learners
July 10, 2026
作者: Letian Wang, Chuhan Zhang, Rishabh Kabra, Jasper Uijlings, Steven Waslander, Andrew Zisserman, Joao Carreira, Kaiming He, Misha Andriluka, Eduard Gabriel Bazavan, Andrei Zanfir, Cristian Sminchisescu
cs.AI
摘要
在下一詞預測的驅動下,自然語言處理從特定任務模型轉變為強大的通用基礎模型。那麼,在電腦視覺中,實現通用模型所需的等價催化劑是什麼?在本文中,我們主張大規模文字轉影片生成可作為電腦視覺中強有力的預訓練範式,為通用視覺智能提供所需的時空先驗、視覺語言對齊和可擴展性。我們提出 GenCeption,該方法利用預訓練的影片生成擴散骨幹網路來定義一個前饋感知模型,能夠根據文字指令執行多種視覺任務。實驗結果表明,GenCeption 在深度估計、表面法向量與相機姿態估計、表情參照分割以及 3D 關鍵點預測等多元任務中達到了最先進的性能,往往能與甚至超越專門模型(例如 DepthAnything3、SAM3、D4RT、VGGT-Omega、Sapiens、David、Genmo 和 Lotus-2)。此外,在可比設定下,影片生成預訓練骨幹網路的表現優於其他預訓練範式(例如 V-JEPA 和 Video MAE)。重要的是,GenCeption 展現了初步的數據與模型擴展特性,以及卓越的數據效率——僅需 D4RT 和 VGGT-Omega 等領先模型 7 到 500 分之一的訓練數據,即可達到與之相當的性能。最後,GenCeption 還展現出引人注目的湧現行為:僅在合成人體影片上訓練的模型,能夠泛化至真實世界影像以及超出分佈的物體類別(例如動物和機器人)。這些發現表明,影片生成不僅僅是一種合成工具,更是邁向物理世界通用視覺智能的基礎路徑。專案頁面:https://genception.github.io
English
Driven by next-token prediction, NLP shifted from task-specific models into powerful generalist foundation models. What, then, is the equivalent catalyst needed to achieve a general-purpose model in computer vision? In this paper, we contend that large-scale text-to-video generation serves as a strong pre-training paradigm for computer vision, providing the necessary spatiotemporal priors, vision-language alignment, and scalability required for general visual intelligence. We introduce GenCeption, which leverages a pre-trained video generative diffusion backbone to define a feed-forward perception model, capable of performing various vision tasks steered by text instructions. Empirical results demonstrate that GenCeption achieves state-of-the-art performance across a diverse suite of tasks, including depth, surface normal, and camera pose estimation, expression-referring segmentation, and 3D keypoint prediction, often matching or surpassing specialized models (e.g. DepthAnything3, SAM3, D4RT, VGGT-Omega, Sapiens, David, Genmo, and Lotus-2). Furthermore, the video generative pretrained backbone outperforms alternative pretraining paradigms (e.g., V-JEPA, and Video MAE) under comparable settings. Importantly, GenCeption exhibits preliminary data and model scaling properties along with exceptional data efficiency, where it achieves comparable performance with leading models like D4RT and VGGT-Omega with 7 to 500 less training data. Finally, GenCeption also exhibits intriguing emergent behaviors: a model trained exclusively on synthetic human videos generalizes to real-world footage and out-of-distribution object categories (e.g., animals and robots). These findings suggest that video generation is not merely a synthesis tool, but a foundational path toward generalist vision intelligence for the physical world. Project page: https://genception.github.io