ChatPaper.aiChatPaper

ビデオからの構造化ダイナミクスの自己教師あり学習

Self-Supervised Learning of Structured Dynamics from Videos

July 23, 2026
著者: Lukas Knobel, Andrew Zisserman, Yuki M. Asano
cs.AI

要旨

映像中の動きの理解は視覚学習における基本的な課題である。なぜならフレーム間の変化には、カメラの動きと物体の動きという二つのダイナミクス源が絡み合っているからである。この分解は表現学習において依然として十分に探求されておらず、その一因はこれらの要因が自然動画において密接に結合しており、個別に教師あり学習することが困難であることにある。しかし、意味のある物体ダイナミクスをカメラ起因の変動から分離したロバストな動き表現を学習するためには、この分解を回復することが重要である。本研究では、事前学習済み画像ビジョントランスフォーマーの凍結特徴量から、このような構造化された動き表現が回復可能かどうかを検証する。我々は構造化ダイナミクスモデル(SDM)を提案する。SDMは、時間的変化の支配的な源と残差ダイナミクスを、単一の絡み合った潜在変数や非構造的で空間的に密な遷移トークンを用いてビデオ変化を表現するのではなく、将来の特徴予測を通じて明示的に分離する。学習は、実動画に対する自己教師あり学習と、合成Kubricデータにおけるシーンダイナミクスの弱教師あり学習を組み合わせる。我々はSDMを、カメラモーション、オブジェクトモーション、およびそれらの組み合わせダイナミクスを含む合成・実動画にわたる新たな評価スイートProbeMotionで評価する。SDMは、グローバルCLSや平均プーリング特徴量を用いたバックボーンベースラインを上回り、いくつかのプローブにおいてはVGGTなどの強教師あり表現と遜色ない性能を示す。これらの結果は、事前学習済み画像モデルが構造化されたビデオダイナミクス表現に容易に転用可能であり、潜在的なビデオダイナミクスの学習と解析に有用な帰納的バイアスを提供することを示唆している。
English
Understanding motion in video is a fundamental challenge for visual learning, as frame-to-frame change entangles two sources of dynamics: camera motion and object motion. This decomposition has remained underexplored in representation learning, partly because these factors are tightly coupled in natural videos and difficult to supervise separately. Yet recovering it is important for learning robust motion representations that separate meaningful object dynamics from camera-induced variation. We study whether such structured motion representations can be recovered from frozen features of a pretrained image vision transformer. We propose the Structured Dynamics Model (SDM), which explicitly separates the dominant source of temporal change from residual dynamics through future-feature prediction, rather than representing video change with a single entangled latent or with unstructured, spatially dense transition tokens. Training combines self-supervised learning on real video with weak supervision of scene dynamics on synthetic Kubric data. We evaluate SDM on ProbeMotion, a new evaluation suite spanning synthetic and real videos with camera motion, object motion, and combined dynamics. SDM outperforms backbone baselines using global CLS or average-pooled features, and compares favorably to strongly supervised representations such as VGGT on several probes, despite using substantially weaker supervision. These results suggest that pretrained image models can be readily repurposed into structured video-dynamics representations, providing a useful inductive bias for learning and analyzing latent video dynamics.