超越形态的运动:基于抽象运动表征的自举式跨类别运动迁移
Motion Beyond Morphology: Bootstrapping Cross-Category Motion Transfer from Abstract Motion Representations
August 3, 2026
作者: Zhixue Fang, Zhimin Zhang, Bi'an Du, Zijie Meng, Yan Zhou, Wei Hu, Guoxin Zhang, Pengfei Wan, Kun Gai
cs.AI
摘要
视频运动迁移旨在利用参考视频中的动态来驱动目标对象。现有方法在很大程度上依赖固定的结构对应,当参考对象与目标对象在形态、关节结构或变形机制上存在显著差异时,这种对应关系便会失去明确定义。我们提出“超越形态的运动迁移”(Motion Beyond Morphology)这一视角,旨在突破固定结构对应的局限,通过保留在不同目标形态下依然具有意义的动态来实现迁移。为实现这一目标,我们提出了一个两阶段框架。第一阶段学习互补的多粒度抽象运动视图,并利用这些视图引导生成跨类别视频对,从而保留跨多样形态的可迁移动态。第二阶段将这种监督内化到直接的参考视频条件生成中,消除了推理时显式运动提取的需要。我们进一步引入了OpenVMT-Dataset和OpenVMT-Bench,用于在相同(Same)、相近(Near)和远距离(Far)类别差距下训练和评估基于图像与文本条件的运动迁移,并计划在论文被接收后发布这两项资源。大量实验表明,本方法在运动保真度和目标保留方面均达到了最先进的性能。项目主页:https://miniz233.github.io/MotionBeyondMorphology/
English
Video motion transfer aims to animate a target object using dynamics from a reference video. Existing formulations largely rely on fixed structural correspondence, which becomes ill-defined when reference and target objects differ substantially in morphology, articulation, or deformation mechanisms. We introduce Motion Beyond Morphology, a perspective that seeks to transfer motion beyond fixed structural correspondence, by preserving dynamics that remain meaningful across different target morphologies. To realize this, we propose a two-stage framework. Stage~I learns complementary multi-granularity abstract motion views and uses them to bootstrap cross-category video pairs that preserve transferable dynamics across diverse morphologies. Stage~II internalizes this supervision into direct reference-video-conditioned generation, removing the need for explicit motion extraction at inference. We further introduce OpenVMT-Dataset and OpenVMT-Bench for training and evaluating image- and text-conditioned motion transfer across Same, Near, and Far category gaps, and plan to release both upon acceptance. Extensive experiments demonstrate state-of-the-art motion fidelity and target preservation. Project page: https://miniz233.github.io/MotionBeyondMorphology/