Hydra-0: 범용 세계 모델링 및 제어를 위한 행동 흐름
Hydra-0: Action Flow for Generalist World Modeling and Control
August 18, 2026
저자: Hongyu Li, Bowen Wen, Xinghao Zhu, Yixuan Wang, Yilun Du, Yunzhu Li, George Konidaris, Stan Birchfield, Soha Pouya, Chenran Li, Yan Chang
cs.AI
초록
우리는 로봇 행동을 픽셀 모션으로 표현하는 행동 흐름(액션 플로우)에 조건화된 범용 세계 모델인 Hydra-0를 소개한다. 이 공유 시각적 인터페이스는 구현체, 작업, 환경, 비디오 생성 백본 전반에 걸쳐 행동 결과를 학습함으로써 범용 세계 모델링과 제어를 가능하게 한다. 최상의 구성은 우리의 행동 조건화 기준선보다 로봇 모션 오차를 90.4%, 객체 모션 오차를 60.2% 낮추면서 제로샷 구성과 데이터 효율적 적응을 지원한다. RoboLab 벤치마크에서 Hydra-0는 재생된 성공률과 참조 성공률 사이의 Pearson 상관계수 r=0.96을 달성한다. 마지막으로, 우리는 이 인터페이스의 창발적 역모드를 발견하는데, 이는 인간 시연에서 전이된 원하는 객체 흐름으로부터 호환되는 로봇 운동을 예측하는 세계 행동 모델이다. 훈련된 행동 헤드는 결과적인 잠재 특징을 실행 가능한 행동으로 매핑하며, 작업별 전문가 로봇 시연을 요구하지 않는다. 종합적으로, 이러한 결과는 이질적인 훈련 데이터, 개루프 정책 평가, 그리고 로봇 제어를 연결하는 공유 제어 인터페이스로서 행동 흐름의 잠재력을 입증한다.
English
We introduce Hydra-0, a generalist world model conditioned on action flow, which represents robot actions as pixel motion. This shared visual interface enables generalist world modeling and control by learning action consequences across embodiments, tasks, environments, and video-generation backbones. Our best configuration achieves 90.4% lower robot-motion error and 60.2% lower object-motion error than our action-conditioned baseline, while supporting zero-shot composition and data-efficient adaptation. On the RoboLab benchmark, Hydra-0 achieves a Pearson correlation of r=0.96 between replayed and reference success rates. Finally, we uncover an emergent inverse mode of this interface: a world action model that predicts compatible robot motion from desired object flow transferred from a human demonstration. A trained action head maps the resulting latent features to executable actions without requiring task-specific expert robot demonstrations. Together, these results demonstrate the potential of action flow as a shared control interface connecting heterogeneous training data, open-loop policy evaluation, and robot control.