비디오로부터 구조적 동역학의 자기 지도 학습
Self-Supervised Learning of Structured Dynamics from Videos
July 23, 2026
저자: Lukas Knobel, Andrew Zisserman, Yuki M. Asano
cs.AI
초록
비디오에서의 움직임 이해는 시각 학습의 근본적인 과제로, 프레임 간 변화는 카메라 움직임과 객체 움직임이라는 두 가지 역학적 원천을 혼합한다. 이러한 분해는 표현 학습에서 충분히 탐구되지 않았으며, 그 이유 중 하나는 이들 요인이 자연 비디오에서 밀접하게 결합되어 있어 개별적으로 감독하기 어렵기 때문이다. 그러나 의미 있는 객체 역학을 카메라 유발 변동과 분리하는 강건한 움직임 표현을 학습하기 위해서는 이러한 복원이 중요하다. 본 연구는 사전 학습된 이미지 비전 트랜스포머의 고정 특징으로부터 이러한 구조화된 움직임 표현을 복원할 수 있는지 조사한다. 우리는 구조화된 역학 모델(Structured Dynamics Model, SDM)을 제안하며, 이는 단일 얽힌 잠재 변수나 구조화되지 않은 공간적으로 조밀한 전이 토큰으로 비디오 변화를 표현하는 대신, 미래 특징 예측을 통해 시간적 변화의 지배적 원천을 잔차 역학으로부터 명시적으로 분리한다. 훈련은 실제 비디오에 대한 자기 지도 학습과 합성 Kubric 데이터의 장면 역학에 대한 약한 감독을 결합한다. 우리는 SDM을 카메라 움직임, 객체 움직임 및 결합 역학을 포함하는 합성 및 실제 비디오에 걸친 새로운 평가 스위트인 ProbeMotion으로 평가한다. SDM은 전역 CLS 또는 평균 풀링 특징을 사용하는 백본 기준선보다 성능이 우수하며, 몇 가지 프로브에서 상당히 약한 감독을 사용함에도 불구하고 VGGT와 같은 강력한 감독 표현에 비해 유리한 성능을 보인다. 이러한 결과는 사전 학습된 이미지 모델이 구조화된 비디오 역학 표현으로 쉽게 활용될 수 있으며, 잠재 비디오 역학의 학습 및 분석에 유용한 귀납적 편향을 제공함을 시사한다.
English
Understanding motion in video is a fundamental challenge for visual learning, as frame-to-frame change entangles two sources of dynamics: camera motion and object motion. This decomposition has remained underexplored in representation learning, partly because these factors are tightly coupled in natural videos and difficult to supervise separately. Yet recovering it is important for learning robust motion representations that separate meaningful object dynamics from camera-induced variation. We study whether such structured motion representations can be recovered from frozen features of a pretrained image vision transformer. We propose the Structured Dynamics Model (SDM), which explicitly separates the dominant source of temporal change from residual dynamics through future-feature prediction, rather than representing video change with a single entangled latent or with unstructured, spatially dense transition tokens. Training combines self-supervised learning on real video with weak supervision of scene dynamics on synthetic Kubric data. We evaluate SDM on ProbeMotion, a new evaluation suite spanning synthetic and real videos with camera motion, object motion, and combined dynamics. SDM outperforms backbone baselines using global CLS or average-pooled features, and compares favorably to strongly supervised representations such as VGGT on several probes, despite using substantially weaker supervision. These results suggest that pretrained image models can be readily repurposed into structured video-dynamics representations, providing a useful inductive bias for learning and analyzing latent video dynamics.