ChatPaper.aiChatPaper

DyPES-VLA: クロス身体性操作のための共有力学事前分布と身体性固有制御の学習

DyPES-VLA: Learning Shared Dynamics Priors and Embodiment-Specific Control for Cross-Embodiment Manipulation

August 6, 2026
著者: Junfeng Li, Junjie He, Zhide Zhong, Yangyang Zheng, Pingyue Sheng, Jiayu Dong, Ruixin Li, Haodong Yan, Jiaguan Zhu, Tianran Zhang, Runze Yu, Wen Chen, Liuqing Yang, Yuxiang Gao, Haoang Li
cs.AI

要旨

視覚・言語・動作(VLA)モデルはロボット操作の強力なパラダイムとなっているが、異種のロボット形態に対する単一の汎用ポリシーを訓練することは依然として未解決の問題である。既存手法には主に2つの限界がある。第一に、多様な視覚データおよび相互作用データに共有されるダイナミクス事前知識を十分に活用しておらず、クロスエンボディメント転送を制限している。第二に、形態固有の行動を共通形式に変換するために大規模な手動前処理を必要とする。これらの限界を克服するため、我々は共有のダイナミクス事前知識と形態固有の制御を学習するクロスエンボディメントVLAであるDyPES-VLAを提案する。まず、クロスエンボディメントデータ上で視覚・言語モデル(VLM)に未来予測を目的とした訓練を行うことで共有のダイナミクス事前知識を学習し、共有クエリ表現が物体の動き、接触、および相互作用に起因するシーン変化を捉えるように導く。次に、形態固有のMixture-of-Experts(MoE)行動ヘッドが、異種の行動を共通形式に手動で事前整列させることなく、これらの共有ダイナミクス事前知識を各形態のネイティブな行動空間で直接実行可能な制御に変換する。このヘッドは、共通の時間的動作構造を捉えるために注意層を共有し、形態固有のフィードフォワード専門家が異なる形態に固有の運動学的制約と制御セマンティクスを解決する。汎用ポリシーとして、提案手法はシミュレーションおよび実世界評価の両方で最先端の性能を達成し、LIBEROで98.0%、RoboCasa-GR1で59.25%、RoboTwin 2.0で89.02%の成功率を記録した。
English
Vision-Language-Action (VLA) models have become a powerful paradigm for robot manipulation, but training a single generalist policy for heterogeneous robot embodiments remains an open problem. Existing methods have two main limitations. First, they underuse dynamics priors shared across diverse visual and interaction data, limiting cross-embodiment transfer. Second, they require extensive manual preprocessing to convert embodiment-specific actions into a common format. To overcome these limitations, we propose DyPES-VLA, a cross-embodiment VLA that learns shared Dynamics Priors and Embodiment-Specific control. First, we learn shared dynamics priors by training the vision-language model (VLM) with a future-prediction objective on cross-embodiment data, driving the shared query representation to capture object motion, contact, and interaction-induced scene changes. Second, an embodiment-specific Mixture-of-Experts (MoE) action head translates these shared dynamics priors into executable controls directly in each embodiment's native action space, without manually pre-aligning heterogeneous actions into a common format. This head shares attention layers to capture common temporal action structures, while its embodiment-specific feed-forward experts resolve the unique kinematic constraints and control semantics of distinct embodiments. As a generalist policy, our \ourmethod achieves state-of-the-art performance across simulation and real-world evaluations, reaching 98.0% success on LIBERO, 59.25% on RoboCasa-GR1, and 89.02% on RoboTwin~2.0.