DyPES-VLA: 교차 신체 조작을 위한 공유 동역학 사전 및 신체별 제어 학습
DyPES-VLA: Learning Shared Dynamics Priors and Embodiment-Specific Control for Cross-Embodiment Manipulation
August 6, 2026
저자: Junfeng Li, Junjie He, Zhide Zhong, Yangyang Zheng, Pingyue Sheng, Jiayu Dong, Ruixin Li, Haodong Yan, Jiaguan Zhu, Tianran Zhang, Runze Yu, Wen Chen, Liuqing Yang, Yuxiang Gao, Haoang Li
cs.AI
초록
비전-언어-행동(Vision-Language-Action, VLA) 모델은 로봇 조작을 위한 강력한 패러다임이 되었지만, 이질적인 로봇 구현(embodiment)을 위한 단일 범용 정책(generalist policy)을 훈련하는 것은 여전히 미해결 문제이다. 기존 방법들은 두 가지 주요 한계를 가진다. 첫째, 다양한 시각 및 상호작용 데이터에 걸쳐 공유되는 동역학 사전(dynamics priors)을 충분히 활용하지 못하여 교차 구현 전이(cross-embodiment transfer)를 제한한다. 둘째, 구현별 행동을 공통 형식으로 변환하기 위해 광범위한 수동 전처리를 요구한다. 이러한 한계를 극복하기 위해, 우리는 공유 동역학 사전(shared Dynamics Priors)과 구현별 제어(Embodiment-Specific control)를 학습하는 교차 구현 VLA인 DyPES-VLA를 제안한다. 첫째, 우리는 교차 구현 데이터에 대한 미래 예측 목적 함수(future-prediction objective)를 사용하여 비전-언어 모델(VLM)을 훈련함으로써 공유 동역학 사전을 학습하며, 이를 통해 공유 쿼리 표현이 객체 운동, 접촉, 상호작용으로 유발된 장면 변화를 포착하도록 한다. 둘째, 구현별 전문가 혼합(Mixture-of-Experts, MoE) 행동 헤드는 이 공유 동역학 사전을 각 구현의 고유 행동 공간(native action space)에서 직접 실행 가능한 제어로 변환하며, 이질적인 행동을 공통 형식으로 수동 사전 정렬할 필요가 없다. 이 헤드는 공통된 시간적 행동 구조를 포착하기 위해 어텐션 계층을 공유하는 반면, 구현별 피드포워드 전문가는 서로 다른 구현의 고유한 운동학적 제약과 제어 의미론을 해결한다. 범용 정책으로서, 우리의 DyPES-VLA는 시뮬레이션 및 실제 환경 평가 전반에서 최첨단 성능을 달성하여 LIBERO에서 98.0%, RoboCasa-GR1에서 59.25%, RoboTwin 2.0에서 89.02%의 성공률을 기록한다.
English
Vision-Language-Action (VLA) models have become a powerful paradigm for robot manipulation, but training a single generalist policy for heterogeneous robot embodiments remains an open problem. Existing methods have two main limitations. First, they underuse dynamics priors shared across diverse visual and interaction data, limiting cross-embodiment transfer. Second, they require extensive manual preprocessing to convert embodiment-specific actions into a common format. To overcome these limitations, we propose DyPES-VLA, a cross-embodiment VLA that learns shared Dynamics Priors and Embodiment-Specific control. First, we learn shared dynamics priors by training the vision-language model (VLM) with a future-prediction objective on cross-embodiment data, driving the shared query representation to capture object motion, contact, and interaction-induced scene changes. Second, an embodiment-specific Mixture-of-Experts (MoE) action head translates these shared dynamics priors into executable controls directly in each embodiment's native action space, without manually pre-aligning heterogeneous actions into a common format. This head shares attention layers to capture common temporal action structures, while its embodiment-specific feed-forward experts resolve the unique kinematic constraints and control semantics of distinct embodiments. As a generalist policy, our \ourmethod achieves state-of-the-art performance across simulation and real-world evaluations, reaching 98.0% success on LIBERO, 59.25% on RoboCasa-GR1, and 89.02% on RoboTwin~2.0.