ShadowDancer:通过从视频及其阴影中学习统一动力学表征,教授视频世界模型任意动作
ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow
July 30, 2026
作者: Jin Cao, Zian Meng, Kaipeng Zhang
cs.AI
摘要
我们提出 ShadowDancer,一种面向交互式视频世界模型的任意动作、帧级控制的新方法。其障碍在于表示层面:现有接口要么对动作进行松散编码,将其如何展开留给模型自行即兴发挥;要么通过结构化信号进行精确编码,但此类信号只能服务某一类动力学且难以获取,因此在不同动力学之间实现精确控制仍然不可行。演示视频是天然的补救方案,可以逐帧指定任意动力学;然而,一段视频仅通过一种特定的外观来展示其动力学,即底层动力学的一个影子,因此从演示中学习到的动作难以迁移到新场景。ShadowDancer 通过两项关键创新解决这一问题:(1)影子对(shadow pairs),即在独立重采样的外观下重放相同动力学的视频对,并由我们的影子库(Shadow Library)大规模构建,从而使得某一动力学族可被精确控制当且仅当能够为其构建这样的视频对;(2)跨影子预测(cross-shadow prediction),即通过从一个影子预测另一个影子来学习动作;因此,配对过程中被重采样的部分在构造上被丢弃,而被保留的部分则成为动作,由此产生统一的动力学表示,驱动一个块因果(block-causal)世界模型。于是,任何演示片段都成为可复用的动作资产,无需动作标签、运动估计器或微调即可在新环境中重放。实验表明,在多样化的动力学族中,相比强潜动作(latent-action)基线和交互式世界模型基线,我们的方法在动作迁移和长时动作展开(rollout)上均有提升,在展开比较中平均盲测胜率达到 86%。视频结果见 https://ShadowDancer-1.github.io。
English
We present ShadowDancer, a novel approach to any-action, frame-level control of interactive video world models. The obstacle is representational: existing interfaces either encode an action loosely, leaving how it unfolds for the model to improvise, or encode it exactly through structured signals that serve one family and are hard to acquire, so precise control across diverse dynamics remains impractical. Demonstration videos are the natural remedy, specifying any dynamics frame by frame; yet a video shows its dynamics only through one particular appearance, a single shadow of the underlying dynamics, so actions learned from demonstrations transfer poorly to new scenes. ShadowDancer addresses this with two key innovations: (1) shadow pairs, video pairs that replay the same dynamics under independently resampled appearance, constructed at scale by our Shadow Library, so that a dynamics family becomes controllable exactly when such pairs can be constructed for it; and (2) cross-shadow prediction, which learns actions by predicting one shadow from the other, so that whatever the pairing resamples is discarded by construction and whatever it preserves becomes the action, yielding a unified dynamics representation that drives a block-causal world model. Any demonstrated clip thus becomes a reusable action asset, replayed in new environments without action labels, motion estimators, or fine-tuning. Experiments demonstrate improved action transfer and long action rollout over strong latent-action and interactive world model baselines across diverse dynamics families, with an average blinded win rate of 86% in rollout comparisons. We show video results at https://ShadowDancer-1.github.io