ChatPaper.aiChatPaper

DyPES-VLA:學習跨具身操作的共享動力學先驗與具身特異性控制

DyPES-VLA: Learning Shared Dynamics Priors and Embodiment-Specific Control for Cross-Embodiment Manipulation

August 6, 2026
作者: Junfeng Li, Junjie He, Zhide Zhong, Yangyang Zheng, Pingyue Sheng, Jiayu Dong, Ruixin Li, Haodong Yan, Jiaguan Zhu, Tianran Zhang, Runze Yu, Wen Chen, Liuqing Yang, Yuxiang Gao, Haoang Li
cs.AI

摘要

視覺-語言-動作(VLA)模型已成為機器人操作的強大範式,但針對異構機器人本體訓練單一通用策略仍是開放性問題。現有方法有兩個主要限制。首先,它們未充分利用跨多樣視覺與互動資料所共享的動力學先驗,限制了跨本體遷移。其次,它們需要大量人工前處理,以將本體特定動作轉換為通用格式。為克服這些限制,我們提出DyPES-VLA,一種學習共享動力學先驗與本體特定控制的跨本體VLA。首先,我們透過在跨本體資料上以未來預測目標訓練視覺-語言模型(VLM),學習共享動力學先驗,促使共享查詢表徵捕捉物體運動、接觸以及互動引起的場景變化。其次,本體特定的專家混合(MoE)動作頭將這些共享動力學先驗直接轉化為每個本體原生動作空間中的可執行控制,無需手動將異構動作預先對齊為通用格式。該動作頭共享注意力層以捕捉常見的時序動作結構,而其本體特定的前饋專家則解析不同本體的獨特運動學約束與控制語意。作為通用策略,我們的DyPES-VLA在模擬與真實世界評估中均達到最佳效能,在LIBERO上達到98.0%成功率,在RoboCasa-GR1上達到59.25%,在RoboTwin 2.0上達到89.02%。
English
Vision-Language-Action (VLA) models have become a powerful paradigm for robot manipulation, but training a single generalist policy for heterogeneous robot embodiments remains an open problem. Existing methods have two main limitations. First, they underuse dynamics priors shared across diverse visual and interaction data, limiting cross-embodiment transfer. Second, they require extensive manual preprocessing to convert embodiment-specific actions into a common format. To overcome these limitations, we propose DyPES-VLA, a cross-embodiment VLA that learns shared Dynamics Priors and Embodiment-Specific control. First, we learn shared dynamics priors by training the vision-language model (VLM) with a future-prediction objective on cross-embodiment data, driving the shared query representation to capture object motion, contact, and interaction-induced scene changes. Second, an embodiment-specific Mixture-of-Experts (MoE) action head translates these shared dynamics priors into executable controls directly in each embodiment's native action space, without manually pre-aligning heterogeneous actions into a common format. This head shares attention layers to capture common temporal action structures, while its embodiment-specific feed-forward experts resolve the unique kinematic constraints and control semantics of distinct embodiments. As a generalist policy, our \ourmethod achieves state-of-the-art performance across simulation and real-world evaluations, reaching 98.0% success on LIBERO, 59.25% on RoboCasa-GR1, and 89.02% on RoboTwin~2.0.