ChatPaper.aiChatPaper

从视频中自监督学习结构化动力学

Self-Supervised Learning of Structured Dynamics from Videos

July 23, 2026
作者: Lukas Knobel, Andrew Zisserman, Yuki M. Asano
cs.AI

摘要

理解视频中的运动是视觉学习的一个基本挑战,因为帧间变化交织着两种动态来源:相机运动和物体运动。这种分解在表示学习中尚未得到充分探索,部分原因是这些因素在自然视频中紧密耦合,难以单独监督。然而,恢复这种分解对于学习稳健的运动表示至关重要,这些表示能够将有意义的物体动态与相机引起的变体分离开来。我们研究这类结构化的运动表示是否可以从预训练图像视觉变换器的冻结特征中恢复。我们提出了结构动力学模型(SDM),该模型通过未来特征预测,明确地将时间变化的主要来源与残余动态分离开来,而不是用单一的纠缠潜在变量或非结构化的、空间密集的转换令牌来表示视频变化。训练结合了在真实视频上的自监督学习以及在合成Kubric数据上对场景动态的弱监督。我们在ProbeMotion上评估SDM,这是一个新的评估套件,涵盖合成和真实视频,包括相机运动、物体运动和组合动态。SDM在多个探针上优于使用全局CLS令牌或平均池化特征的主干基线,并且与强监督表示(如VGGT)相比表现相当,尽管使用的监督更弱。这些结果表明,预训练图像模型可以很容易地被重新用于结构化的视频动态表示,为学习和分析潜在视频动态提供了有用的归纳偏置。
English
Understanding motion in video is a fundamental challenge for visual learning, as frame-to-frame change entangles two sources of dynamics: camera motion and object motion. This decomposition has remained underexplored in representation learning, partly because these factors are tightly coupled in natural videos and difficult to supervise separately. Yet recovering it is important for learning robust motion representations that separate meaningful object dynamics from camera-induced variation. We study whether such structured motion representations can be recovered from frozen features of a pretrained image vision transformer. We propose the Structured Dynamics Model (SDM), which explicitly separates the dominant source of temporal change from residual dynamics through future-feature prediction, rather than representing video change with a single entangled latent or with unstructured, spatially dense transition tokens. Training combines self-supervised learning on real video with weak supervision of scene dynamics on synthetic Kubric data. We evaluate SDM on ProbeMotion, a new evaluation suite spanning synthetic and real videos with camera motion, object motion, and combined dynamics. SDM outperforms backbone baselines using global CLS or average-pooled features, and compares favorably to strongly supervised representations such as VGGT on several probes, despite using substantially weaker supervision. These results suggest that pretrained image models can be readily repurposed into structured video-dynamics representations, providing a useful inductive bias for learning and analyzing latent video dynamics.