DyPES-VLA:面向跨具身操作的共享动态先验学习与具身特定控制
DyPES-VLA: Learning Shared Dynamics Priors and Embodiment-Specific Control for Cross-Embodiment Manipulation
August 6, 2026
作者: Junfeng Li, Junjie He, Zhide Zhong, Yangyang Zheng, Pingyue Sheng, Jiayu Dong, Ruixin Li, Haodong Yan, Jiaguan Zhu, Tianran Zhang, Runze Yu, Wen Chen, Liuqing Yang, Yuxiang Gao, Haoang Li
cs.AI
摘要
视觉-语言-动作(VLA)模型已成为机器人操作领域的一种强大范式,但为异构机器人本体训练单一通用策略仍是一个开放性问题。现有方法存在两个主要局限。首先,它们未能充分利用跨多样视觉与交互数据共享的动力学先验,限制了跨本体迁移能力。其次,它们需要大量人工预处理以将本体特有的动作转换为通用格式。为克服上述局限,我们提出DyPES-VLA,一种学习共享动力学先验(Dynamics Priors)与本体特有控制(Embodiment-Specific control)的跨本体VLA模型。首先,我们通过在跨本体数据上以未来预测目标训练视觉-语言模型(VLM)来学习共享动力学先验,促使共享查询表示捕捉物体运动、接触以及交互引发的场景变化。其次,一个本体特有的专家混合(MoE)动作头将这些共享动力学先验直接转换为各本体原生动作空间中可执行的控制指令,无需将异构动作人工预对齐为通用格式。该动作头共享注意力层以捕获共有的时序动作结构,同时其本体特有的前馈专家负责解析不同本体独特的运动学约束与控制语义。作为通用策略,我们的方法在仿真与真实世界评估中均达到最先进性能,在LIBERO上取得98.0%的成功率,在RoboCasa-GR1上达到59.25%,在RoboTwin 2.0上达到89.02%。
English
Vision-Language-Action (VLA) models have become a powerful paradigm for robot manipulation, but training a single generalist policy for heterogeneous robot embodiments remains an open problem. Existing methods have two main limitations. First, they underuse dynamics priors shared across diverse visual and interaction data, limiting cross-embodiment transfer. Second, they require extensive manual preprocessing to convert embodiment-specific actions into a common format. To overcome these limitations, we propose DyPES-VLA, a cross-embodiment VLA that learns shared Dynamics Priors and Embodiment-Specific control. First, we learn shared dynamics priors by training the vision-language model (VLM) with a future-prediction objective on cross-embodiment data, driving the shared query representation to capture object motion, contact, and interaction-induced scene changes. Second, an embodiment-specific Mixture-of-Experts (MoE) action head translates these shared dynamics priors into executable controls directly in each embodiment's native action space, without manually pre-aligning heterogeneous actions into a common format. This head shares attention layers to capture common temporal action structures, while its embodiment-specific feed-forward experts resolve the unique kinematic constraints and control semantics of distinct embodiments. As a generalist policy, our \ourmethod achieves state-of-the-art performance across simulation and real-world evaluations, reaching 98.0% success on LIBERO, 59.25% on RoboCasa-GR1, and 89.02% on RoboTwin~2.0.