ChatPaper.aiChatPaper

超越形態的動作:從抽象動作表徵引導跨類別動作遷移

Motion Beyond Morphology: Bootstrapping Cross-Category Motion Transfer from Abstract Motion Representations

August 3, 2026
作者: Zhixue Fang, Zhimin Zhang, Bi'an Du, Zijie Meng, Yan Zhou, Wei Hu, Guoxin Zhang, Pengfei Wan, Kun Gai
cs.AI

摘要

視頻運動遷移旨在利用參考視頻的動態來驅動目標對象。現有方法大多依賴固定的結構對應,但當參考對象與目標對象在形態、關節結構或形變機制上差異顯著時,這種對應關係便難以定義。我們提出「超越形態的運動」(Motion Beyond Morphology)這一視角,旨在超越固定結構對應來遷移運動,保留在不同目標形態下仍具意義的動態。為實現此目標,我們提出一個兩階段框架。第一階段學習互補的多粒度抽象運動視角,並利用這些視角引導生成跨類別視頻對,以保留在多樣形態間可遷移的動態。第二階段將此監督信號內化為直接以參考視頻為條件的生成過程,從而無需在推理時進行顯式運動提取。我們進一步推出 OpenVMT-Dataset 與 OpenVMT-Bench,用於訓練和評估在「同類」、「近類」與「遠類」類別差距下的圖像與文本條件運動遷移,並計劃在論文接收後公開釋出。大量實驗結果顯示,我們的方法在運動保真度與目標保持方面達到了當前最佳水平。項目主頁:https://miniz233.github.io/MotionBeyondMorphology/
English
Video motion transfer aims to animate a target object using dynamics from a reference video. Existing formulations largely rely on fixed structural correspondence, which becomes ill-defined when reference and target objects differ substantially in morphology, articulation, or deformation mechanisms. We introduce Motion Beyond Morphology, a perspective that seeks to transfer motion beyond fixed structural correspondence, by preserving dynamics that remain meaningful across different target morphologies. To realize this, we propose a two-stage framework. Stage~I learns complementary multi-granularity abstract motion views and uses them to bootstrap cross-category video pairs that preserve transferable dynamics across diverse morphologies. Stage~II internalizes this supervision into direct reference-video-conditioned generation, removing the need for explicit motion extraction at inference. We further introduce OpenVMT-Dataset and OpenVMT-Bench for training and evaluating image- and text-conditioned motion transfer across Same, Near, and Far category gaps, and plan to release both upon acceptance. Extensive experiments demonstrate state-of-the-art motion fidelity and target preservation. Project page: https://miniz233.github.io/MotionBeyondMorphology/