Hydra-0:通用世界建模與控制的動作流
Hydra-0: Action Flow for Generalist World Modeling and Control
August 18, 2026
作者: Hongyu Li, Bowen Wen, Xinghao Zhu, Yixuan Wang, Yilun Du, Yunzhu Li, George Konidaris, Stan Birchfield, Soha Pouya, Chenran Li, Yan Chang
cs.AI
摘要
我們介紹Hydra-0,一種以動作流為條件的通用世界模型,將機器人動作表示為像素運動。此共享視覺介面透過學習跨具體形態、任務、環境及影片生成骨幹的動作後果,實現通用的世界建模與控制。我們的最佳配置在機器人運動誤差上較動作條件基準降低90.4%,在物體運動誤差上降低60.2%,同時支援零樣本組合與資料高效率的適應。在RoboLab基準上,Hydra-0在重播與參考成功率之間達到皮爾遜相關係數r=0.96。最後,我們揭示了此介面的一種湧現逆向模式:一種世界動作模型,可從人類示範轉移的期望物體流預測相容的機器人運動。經過訓練的動作頭可將產生的潛在特徵映射為可執行動作,無需特定任務的專家機器人示範。綜合以上結果,這些成果展示了動作流作為一種共享控制介面的潛力,可連結異質訓練資料、開環策略評估與機器人控制。
English
We introduce Hydra-0, a generalist world model conditioned on action flow, which represents robot actions as pixel motion. This shared visual interface enables generalist world modeling and control by learning action consequences across embodiments, tasks, environments, and video-generation backbones. Our best configuration achieves 90.4% lower robot-motion error and 60.2% lower object-motion error than our action-conditioned baseline, while supporting zero-shot composition and data-efficient adaptation. On the RoboLab benchmark, Hydra-0 achieves a Pearson correlation of r=0.96 between replayed and reference success rates. Finally, we uncover an emergent inverse mode of this interface: a world action model that predicts compatible robot motion from desired object flow transferred from a human demonstration. A trained action head maps the resulting latent features to executable actions without requiring task-specific expert robot demonstrations. Together, these results demonstrate the potential of action flow as a shared control interface connecting heterogeneous training data, open-loop policy evaluation, and robot control.