ChatPaper.aiChatPaper

從影片中進行結構化動態的自監督學習

Self-Supervised Learning of Structured Dynamics from Videos

July 23, 2026
作者: Lukas Knobel, Andrew Zisserman, Yuki M. Asano
cs.AI

摘要

理解影片中的運動是視覺學習的一項基本挑戰,因為影格之間的變化糾纏了兩種動態來源:攝影機運動與物體運動。這種分解在表徵學習中仍未被充分探索,部分原因在於這些因素在自然影片中緊密耦合,難以分別進行監督。然而,恢復這種分解對於學習能夠將有意義的物體動態與攝影機引起的變化區分開來的穩健運動表徵至關重要。我們研究是否可以從預訓練的影像視覺變換器的凍結特徵中恢復此類結構化運動表徵。我們提出了結構化動態模型(Structured Dynamics Model, SDM),該模型通過未來特徵預測,明確地將時間變化的主要來源與殘餘動態分離開來,而不是使用單一糾纏的潛在變量或非結構化的空間密集轉換標記來表示影片變化。訓練結合了在真實影片上的自監督式學習與在合成Kubric數據上對場景動態的弱監督。我們在ProbeMotion上評估SDM,這是一個新的評估套件,涵蓋了具有攝影機運動、物體運動及兩者結合動態的合成與真實影片。SDM在使用全域CLS或平均池化特徵的骨幹基線方法中表現更優,並且在若干探測任務上與強監督表徵(如VGGT)相比毫不遜色,儘管使用了明顯較弱的監督信號。這些結果表明,預訓練的影像模型可以輕易地被重新用於結構化的影片動態表徵,為學習與分析潛在影片動態提供了有用的歸納偏置。
English
Understanding motion in video is a fundamental challenge for visual learning, as frame-to-frame change entangles two sources of dynamics: camera motion and object motion. This decomposition has remained underexplored in representation learning, partly because these factors are tightly coupled in natural videos and difficult to supervise separately. Yet recovering it is important for learning robust motion representations that separate meaningful object dynamics from camera-induced variation. We study whether such structured motion representations can be recovered from frozen features of a pretrained image vision transformer. We propose the Structured Dynamics Model (SDM), which explicitly separates the dominant source of temporal change from residual dynamics through future-feature prediction, rather than representing video change with a single entangled latent or with unstructured, spatially dense transition tokens. Training combines self-supervised learning on real video with weak supervision of scene dynamics on synthetic Kubric data. We evaluate SDM on ProbeMotion, a new evaluation suite spanning synthetic and real videos with camera motion, object motion, and combined dynamics. SDM outperforms backbone baselines using global CLS or average-pooled features, and compares favorably to strongly supervised representations such as VGGT on several probes, despite using substantially weaker supervision. These results suggest that pretrained image models can be readily repurposed into structured video-dynamics representations, providing a useful inductive bias for learning and analyzing latent video dynamics.