ShadowDancer:ビデオとその影から統一的ダイナミクス表現を学習することで、ビデオ世界モデルに任意のアクションを教示する
ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow
July 30, 2026
著者: Jin Cao, Zian Meng, Kaipeng Zhang
cs.AI
要旨
本稿では、インタラクティブビデオ世界モデルにおける任意アクションのフレームレベル制御を実現する新しい手法、ShadowDancerを提案する。その障害は表現上のものにある。既存のインターフェースは、アクションを緩く符号化してその展開をモデルの即興に委ねるか、あるいは特定のファミリーにしか対応せず取得が困難な構造化信号を通じてアクションを厳密に符号化するかのどちらかであり、多様なダイナミクスにわたる精密な制御は依然として実用的でない。デモンストレーションビデオは、任意のダイナミクスをフレーム単位で指定できる自然な解決策である。しかし、ビデオはそのダイナミクスを特定の見た目を通して、すなわち根底にあるダイナミクスの単一の影としてしか示さないため、デモンストレーションから学習されたアクションは新しいシーンへの転移が乏しい。ShadowDancerは、以下の2つの主要な革新によってこの問題に対処する。(1) シャドウペア:独立に再サンプリングされた見た目の下で同一のダイナミクスを再生するビデオペアであり、Shadow Libraryによって大規模に構築される。これにより、あるダイナミクスファミリーは、そのようなペアが構築可能である場合に限り制御可能になる。(2) クロスシャドウ予測:一方のシャドウからもう一方を予測することでアクションを学習する。これにより、ペアリングが再サンプリングするものは設計上破棄され、保存されるものがアクションとなる。その結果、ブロック因果的世界モデルを駆動する統一的なダイナミクス表現が得られる。したがって、デモンストレーションされた任意のクリップは再利用可能なアクション資産となり、アクションラベル、モーション推定器、ファインチューニングを必要とせずに新しい環境で再生できる。実験では、多様なダイナミクスファミリーにわたって、強力な潜在アクションベースラインおよびインタラクティブ世界モデルベースラインと比較し、アクション転移と長期的なアクションロールアウトの改善を示した。ロールアウト比較における平均ブラインド勝率は86%である。ビデオ結果は https://ShadowDancer-1.github.io で公開している。
English
We present ShadowDancer, a novel approach to any-action, frame-level control of interactive video world models. The obstacle is representational: existing interfaces either encode an action loosely, leaving how it unfolds for the model to improvise, or encode it exactly through structured signals that serve one family and are hard to acquire, so precise control across diverse dynamics remains impractical. Demonstration videos are the natural remedy, specifying any dynamics frame by frame; yet a video shows its dynamics only through one particular appearance, a single shadow of the underlying dynamics, so actions learned from demonstrations transfer poorly to new scenes. ShadowDancer addresses this with two key innovations: (1) shadow pairs, video pairs that replay the same dynamics under independently resampled appearance, constructed at scale by our Shadow Library, so that a dynamics family becomes controllable exactly when such pairs can be constructed for it; and (2) cross-shadow prediction, which learns actions by predicting one shadow from the other, so that whatever the pairing resamples is discarded by construction and whatever it preserves becomes the action, yielding a unified dynamics representation that drives a block-causal world model. Any demonstrated clip thus becomes a reusable action asset, replayed in new environments without action labels, motion estimators, or fine-tuning. Experiments demonstrate improved action transfer and long action rollout over strong latent-action and interactive world model baselines across diverse dynamics families, with an average blinded win rate of 86% in rollout comparisons. We show video results at https://ShadowDancer-1.github.io