Hydra-0:汎用的世界モデリングと制御のための行動フロー
Hydra-0: Action Flow for Generalist World Modeling and Control
August 18, 2026
著者: Hongyu Li, Bowen Wen, Xinghao Zhu, Yixuan Wang, Yilun Du, Yunzhu Li, George Konidaris, Stan Birchfield, Soha Pouya, Chenran Li, Yan Chang
cs.AI
要旨
本稿では、アクションフローに条件付けられた汎用世界モデルHydra-0を紹介する。Hydra-0はロボットの動作をピクセル運動として表現する。この共有視覚インターフェースにより、ロボット形態、タスク、環境、ビデオ生成バックボーンにわたって動作の結果を学習することで、汎用的な世界モデリングと制御が可能になる。最良構成では、アクション条件付きベースラインと比較して、ロボット運動誤差を90.4%、物体運動誤差を60.2%低減し、同時にゼロショット合成とデータ効率的な適応を実現する。RoboLabベンチマークでは、Hydra-0は再生成功率と参照成功率の間でピアソン相関 r=0.96 を達成する。最後に、このインターフェースの創発的な逆モードを明らかにする。すなわち、人間のデモンストレーションから転移された所望の物体フローから整合的なロボット運動を予測する世界行動モデルである。訓練されたアクションヘッドは、得られた潜在特徴を、タスク固有の専門家ロボットデモンストレーションを必要とせずに実行可能な動作へとマッピングする。これらの結果は総合すると、アクションフローが、異種の学習データ、オープンループ方策評価、ロボット制御を結びつける共有制御インターフェースとしての可能性を示している。
English
We introduce Hydra-0, a generalist world model conditioned on action flow, which represents robot actions as pixel motion. This shared visual interface enables generalist world modeling and control by learning action consequences across embodiments, tasks, environments, and video-generation backbones. Our best configuration achieves 90.4% lower robot-motion error and 60.2% lower object-motion error than our action-conditioned baseline, while supporting zero-shot composition and data-efficient adaptation. On the RoboLab benchmark, Hydra-0 achieves a Pearson correlation of r=0.96 between replayed and reference success rates. Finally, we uncover an emergent inverse mode of this interface: a world action model that predicts compatible robot motion from desired object flow transferred from a human demonstration. A trained action head maps the resulting latent features to executable actions without requiring task-specific expert robot demonstrations. Together, these results demonstrate the potential of action flow as a shared control interface connecting heterogeneous training data, open-loop policy evaluation, and robot control.