ChatPaper.aiChatPaper

ShadowDancer:通過從視頻及其影子中學習統一的動力學表徵,教導視頻世界模型執行任意動作

ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow

July 30, 2026
作者: Jin Cao, Zian Meng, Kaipeng Zhang
cs.AI

摘要

我們提出 ShadowDancer,一種對互動式影片世界模型進行任意動作、逐幀控制的新方法。其障礙在於表徵層面:現有介面要麼鬆散地編碼動作,讓模型自行即興演繹其展開方式;要麼透過結構化信號精確編碼動作,但這些信號僅服務於單一動力學家族且難以獲取,因此跨多樣動力學的精確控制仍不切實際。示範影片是自然的解決方案,能逐幀指定任意動力學;然而,影片僅透過某一特定外觀來展示其動力學,也就是底層動力學的單一影子,因此從示範中學到的動作難以遷移到新場景。ShadowDancer 以兩項關鍵創新解決此問題:(1) 影子對——在獨立重新採樣的外觀下重播相同動力學的影片對,由我們的 Shadow Library 大規模構建;因此,一個動力學家族唯有在能為其構建此類對時,才得以被精確控制;(2) 跨影子預測——透過從一個影子預測另一個影子來學習動作;如此一來,配對所重新採樣的一切在構造上被丟棄,而其所保留的一切則成為動作,進而產生統一的動力學表徵,驅動區塊因果世界模型。因此,任何示範片段都成為可重複使用的動作資產,可在無動作標籤、無運動估計器、也無需微調的情況下,在新環境中重播。實驗表明,在多樣動力學家族中,相較於強大的潛在動作基線與互動式世界模型基線,本方法在動作遷移與長程動作展開上均有改進,在展開比較中的平均盲測勝率為 86%。我們在 https://ShadowDancer-1.github.io 展示影片結果。
English
We present ShadowDancer, a novel approach to any-action, frame-level control of interactive video world models. The obstacle is representational: existing interfaces either encode an action loosely, leaving how it unfolds for the model to improvise, or encode it exactly through structured signals that serve one family and are hard to acquire, so precise control across diverse dynamics remains impractical. Demonstration videos are the natural remedy, specifying any dynamics frame by frame; yet a video shows its dynamics only through one particular appearance, a single shadow of the underlying dynamics, so actions learned from demonstrations transfer poorly to new scenes. ShadowDancer addresses this with two key innovations: (1) shadow pairs, video pairs that replay the same dynamics under independently resampled appearance, constructed at scale by our Shadow Library, so that a dynamics family becomes controllable exactly when such pairs can be constructed for it; and (2) cross-shadow prediction, which learns actions by predicting one shadow from the other, so that whatever the pairing resamples is discarded by construction and whatever it preserves becomes the action, yielding a unified dynamics representation that drives a block-causal world model. Any demonstrated clip thus becomes a reusable action asset, replayed in new environments without action labels, motion estimators, or fine-tuning. Experiments demonstrate improved action transfer and long action rollout over strong latent-action and interactive world model baselines across diverse dynamics families, with an average blinded win rate of 86% in rollout comparisons. We show video results at https://ShadowDancer-1.github.io