비디오 생성 모델은 범용 비전 학습기이다
Video Generation Models are General-Purpose Vision Learners
July 10, 2026
저자: Letian Wang, Chuhan Zhang, Rishabh Kabra, Jasper Uijlings, Steven Waslander, Andrew Zisserman, Joao Carreira, Kaiming He, Misha Andriluka, Eduard Gabriel Bazavan, Andrei Zanfir, Cristian Sminchisescu
cs.AI
초록
다음 토큰 예측에 의해 주도되면서, NLP는 작업별 모델에서 강력한 범용 기반 모델로 전환되었습니다. 그렇다면 컴퓨터 비전에서 범용 모델을 달성하기 위해 필요한 동등한 촉매는 무엇일까요? 본 논문에서 우리는 대규모 텍스트-비디오 생성이 컴퓨터 비전을 위한 강력한 사전 학습 패러다임으로 작용하여, 일반 시각 지능에 필요한 시공간적 사전 정보, 시각-언어 정렬 및 확장성을 제공한다고 주장합니다. 우리는 GenCeption을 소개합니다. 이는 사전 학습된 비디오 생성 확산 백본을 활용하여 피드포워드 인식 모델을 정의하며, 텍스트 명령에 따라 다양한 비전 작업을 수행할 수 있습니다. 실험 결과는 GenCeption이 깊이, 표면 법선, 카메라 포즈 추정, 표정 기반 분할, 3D 키포인트 예측 등 다양한 작업에서 최첨단 성능을 달성하며, 종종 전문 모델(예: DepthAnything3, SAM3, D4RT, VGGT-Omega, Sapiens, David, Genmo, Lotus-2)과 일치하거나 능가함을 보여줍니다. 또한 비디오 생성 사전 학습 백본은 유사한 설정에서 대체 사전 학습 패러다임(예: V-JEPA, Video MAE)보다 우수한 성능을 보입니다. 중요하게도, GenCeption은 초기 데이터 및 모델 확장 속성과 함께 뛰어난 데이터 효율성을 나타내며, 7~500배 적은 훈련 데이터로 D4RT 및 VGGT-Omega와 같은 선도적 모델과 비슷한 성능을 달성합니다. 마지막으로 GenCeption은 흥미로운 창발적 행동을 보여줍니다: 합성 인간 비디오만으로 훈련된 모델이 실제 영상과 분포 외 객체 범주(예: 동물 및 로봇)에 일반화됩니다. 이러한 결과는 비디오 생성이 단순한 합성 도구가 아니라 물리적 세계를 위한 범용 비전 지능으로 가는 기초적인 경로임을 시사합니다. 프로젝트 페이지: https://genception.github.io
English
Driven by next-token prediction, NLP shifted from task-specific models into powerful generalist foundation models. What, then, is the equivalent catalyst needed to achieve a general-purpose model in computer vision? In this paper, we contend that large-scale text-to-video generation serves as a strong pre-training paradigm for computer vision, providing the necessary spatiotemporal priors, vision-language alignment, and scalability required for general visual intelligence. We introduce GenCeption, which leverages a pre-trained video generative diffusion backbone to define a feed-forward perception model, capable of performing various vision tasks steered by text instructions. Empirical results demonstrate that GenCeption achieves state-of-the-art performance across a diverse suite of tasks, including depth, surface normal, and camera pose estimation, expression-referring segmentation, and 3D keypoint prediction, often matching or surpassing specialized models (e.g. DepthAnything3, SAM3, D4RT, VGGT-Omega, Sapiens, David, Genmo, and Lotus-2). Furthermore, the video generative pretrained backbone outperforms alternative pretraining paradigms (e.g., V-JEPA, and Video MAE) under comparable settings. Importantly, GenCeption exhibits preliminary data and model scaling properties along with exceptional data efficiency, where it achieves comparable performance with leading models like D4RT and VGGT-Omega with 7 to 500 less training data. Finally, GenCeption also exhibits intriguing emergent behaviors: a model trained exclusively on synthetic human videos generalizes to real-world footage and out-of-distribution object categories (e.g., animals and robots). These findings suggest that video generation is not merely a synthesis tool, but a foundational path toward generalist vision intelligence for the physical world. Project page: https://genception.github.io