ShadowDancer: 비디오와 그 그림자로부터 통합된 역학 표현을 학습하여 비디오 월드 모델에 임의의 행동을 가르치는 방법
ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow
July 30, 2026
저자: Jin Cao, Zian Meng, Kaipeng Zhang
cs.AI
초록
우리는 인터랙티브 비디오 월드 모델에서 임의의 동작을 프레임 수준으로 제어하는 새로운 접근 방식인 ShadowDancer를 제안한다. 장애물은 표현(representation)에 있다. 기존 인터페이스는 동작을 느슨하게 인코딩하여 그것이 어떻게 전개될지를 모델이 즉흥적으로 처리하도록 맡기거나, 하나의 계열에만 적용되고 획득하기 어려운 구조화된 신호를 통해 동작을 정확히 인코딩한다. 따라서 다양한 역학에 걸친 정밀한 제어는 여전히 비실용적이다. 데모 비디오는 임의의 역학을 프레임 단위로 명시하는 자연스러운 해결책이지만, 비디오는 하나의 특정한 외관, 즉 기본 역학의 단일 그림자를 통해서만 자신의 역학을 보여주므로, 데모에서 학습된 동작은 새로운 장면으로 잘 전이되지 않는다. ShadowDancer는 두 가지 핵심 혁신을 통해 이 문제를 해결한다. (1) 그림자 쌍(shadow pair)은 동일한 역학을 독립적으로 재표본화된 외관으로 재생하는 비디오 쌍으로, Shadow Library가 대규모로 구축하며, 어떤 역학 계열은 그러한 쌍을 구축할 수 있을 때에만 정확히 제어 가능해진다. (2) 교차 그림자 예측(cross-shadow prediction)은 한 그림자로부터 다른 그림자를 예측하여 동작을 학습하며, 쌍이 재표본화하는 것은 설계상 폐기되고 보존하는 것이 동작이 된다. 이를 통해 블록-인과적 월드 모델을 구동하는 통합된 역학 표현이 생성된다. 따라서 데모된 모든 클립은 동작 레이블, 모션 추정기, 미세 조정 없이도 새로운 환경에서 재생할 수 있는 재사용 가능한 동작 자산이 된다. 실험은 다양한 역학 계열에 걸쳐 강력한 잠재 동작 및 인터랙티브 월드 모델 베이스라인보다 향상된 동작 전이와 긴 동작 롤아웃을 보여주며, 롤아웃 비교에서 평균 블라인드 승률 86%를 기록했다. 비디오 결과는 https://ShadowDancer-1.github.io 에서 확인할 수 있다.
English
We present ShadowDancer, a novel approach to any-action, frame-level control of interactive video world models. The obstacle is representational: existing interfaces either encode an action loosely, leaving how it unfolds for the model to improvise, or encode it exactly through structured signals that serve one family and are hard to acquire, so precise control across diverse dynamics remains impractical. Demonstration videos are the natural remedy, specifying any dynamics frame by frame; yet a video shows its dynamics only through one particular appearance, a single shadow of the underlying dynamics, so actions learned from demonstrations transfer poorly to new scenes. ShadowDancer addresses this with two key innovations: (1) shadow pairs, video pairs that replay the same dynamics under independently resampled appearance, constructed at scale by our Shadow Library, so that a dynamics family becomes controllable exactly when such pairs can be constructed for it; and (2) cross-shadow prediction, which learns actions by predicting one shadow from the other, so that whatever the pairing resamples is discarded by construction and whatever it preserves becomes the action, yielding a unified dynamics representation that drives a block-causal world model. Any demonstrated clip thus becomes a reusable action asset, replayed in new environments without action labels, motion estimators, or fine-tuning. Experiments demonstrate improved action transfer and long action rollout over strong latent-action and interactive world model baselines across diverse dynamics families, with an average blinded win rate of 86% in rollout comparisons. We show video results at https://ShadowDancer-1.github.io