ChatPaper.aiChatPaper

기하학 학습자로서의 비디오 생성 모델

Video Generative Models as Geometry Learner

August 28, 2026
저자: Haosen Yang, Jifei Song, Zhensong Zhang, Xiatian Zhu, Jiankang Deng
cs.AI

초록

최근 기하 추정을 위한 생성적 접근법들은 사전 학습된 이미지 확산 모델을 적용하여 이 작업을 이미지 조건 생성으로 취급한다. 기성 이미지 확산 모델을 활용하는 이러한 접근법들은 (i) 깊이 및 표면 법선 추정을 위한 작업별 기하 모델을 독립적으로 훈련시켜 이러한 기하 대상들의 내재적 상관관계를 탐구할 기회를 잃거나, (ii) 수정된 이미지 확산 백본(예: 변경된 자기 주의 메커니즘)을 공동으로 미세 조정하는데, 이는 일반적으로 상당한 양의 레이블된 데이터를 요구한다. 이러한 한계를 원칙적인 방식으로 극복하기 위해, 우리는 사전 학습된 비디오 생성 모델을 기하 추정을 위한 통합적이고 데이터 효율적인 프레임워크로 재목적화하며, 이를 혁신적으로 다음 프레임 예측 작업으로 공식화한다. 우리의 방법인 GeoNeXt는 비디오 모델로부터 자연스럽게 구조화된 지식과 더 풍부한 사전 정보를 상속받으며, 이를 이미지와 기하 대상의 공동 모델링(이미지 <-> 기하)에 적응시켜 기하 학습을 더욱 데이터 효율적이고 효과적으로 만든다. 광범위한 실험을 통해 다양한 데이터셋에서 제로샷 단안 깊이 및 표면 법선 추정에 대한 우리의 방법을 검증했으며, 훨씬 적은 훈련 데이터를 사용하면서도 이전의 작업별 및 통합 생성 경쟁 방법들을 능가한다. 특히, 우리의 방법은 100배 이상 많은 데이터로 훈련된 판별적 최신 기법들과 견줄 만하며, 여러 벤치마크에서 두드러진 성능을 보인다.
English
Recent generative approaches to geometry estimation adapt pretrained image diffusion models and treat the task as image-conditioned generation. Leveraging off-the-shelf image diffusion models, they either (i) train task-specific geometry models (for depth and surface normal estimation) independently, losing the opportunity of exploring the intrinsic correlation of these geometric targets, or (ii) jointly fine-tune modified image diffusion backbones (e.g., altered self-attention), which typically demands substantial labeled data. To overcome these limitations in a principled fashion, we repurpose pretrained video generative models as a unified and data-efficient framework for geometry estimation, formulated innovatively as a next-frames prediction task. Our method, GeoNeXt, inherits naturally structured knowledge and richer priors from the video model, while further adapting them for joint modeling of images and geometry targets (image <-> geometry), enabling more data efficient and effective learning of geometry. Extensive experiments validate our method for zero-shot monocular depth and surface normal estimation across diverse datasets, outperforming both previous task-specific and unified generative competitors while using substantially less training data. Notably, our method rivals discriminative state-of-the-art approaches trained on over 100x more data and even standouts on several benchmarks.