GE-Act 2.0:預訓練與擴展用於機器人操作的世界-動作模型
GE-Act 2.0: Pretraining and Scaling a World-Action Model for Robotic Manipulation
September 4, 2026
作者: AgiBot Research Team, Renhang Liu, Wenzhi Zhao, Zhuo Yang, Liliang Chen, Pengfei Zhou, Shengcong Chen, Guanghui Ren, Youlun Peng, Rongjun Jin, Nan Wang, Sukai Wang, Xindong He, Jinyuan Feng, Ziyu Xiong, Linqing Zhong, Yifei Wei, Feng Han, Long Zhang, Da Huang, Nanshu Zhao, Chenghao Yin, Mo Wu, Zhaodong Yan, Kongtao Hu, Yuxiang Yan, Aogelijiang Niyazi, Yu Fang, Jia Zeng, Lizhu Meng, Daizhen Lv, Haoyu Cao, Zhiwen Hou, Lianjin Ye, Yuehan Niu, Zhikai Cai, Xuan Hu, Hui Min, Xiongfeng Cai, Yue Liao, Jing Wu, Soujanya Poria, Ye Li, Sanping Zhou, Maoqing Yao
cs.AI
摘要
世界行動模型 (WAM) 預測未來狀態以引導機器人動作,從而能夠從無動作影片和帶動作標籤的互動中學習。大多數模型繼承了預訓練的影片生成器,使得 WAM 的預訓練和規模化探索不足。我們引入了 Genie Envisioner Act 2.0 (GE-Act 2.0),這是一個世界行動模型,其可訓練的生成和動作組件全部在操作資料上從頭初始化。它結合了控制導向自編碼器 (CoAE)、單步視覺規劃器 (SVP) 和逆動力學模型 (IDM)。CoAE 在激進壓縮下保留與動作和指令相關的資訊,而 SVP 在一個可微分步驟中產生完整的未來狀態,因此視覺規劃和逆動力學可以分別在互補資料上進行預訓練。然後,這些組件使用知識對齊選擇性優化 (KASO) 進行聯合訓練,該方法透過僅選擇被判斷為與記錄動作行為相容的預測未來,來減少不匹配的監督。我們直接評估預訓練檢查點,無需針對每個任務進行微調,在 20 個操作技能組的 100 個任務上,使用保留的場景、背景、光照和物體實例。將協同訓練資料從 300 小時擴展到 30,000 小時,將 G1-OP 上的成功率從 17.1% 提高到 44.1%,並將 G2-90D 上的成功率從 13.4% 提高到 31.1%;儘管 G2-90D 佔協同訓練資料不到 2%,但其提升了 17.7 個百分點,顯示出跨具身轉移。增益涵蓋 19/20 和 18/20 個技能組,且技能特定覆蓋率與零樣本分布外 (OOD) 成功率高度相關(皮爾森 r=0.80;斯皮爾曼 rho=0.85)。在相同協議下,該模型在至少 90% 的試驗中能錨定物體、顏色、形狀和位置的指稱,並且即使明確指令與已承諾的行為或常規場景關聯相衝突,也能遵循這些指令。
English
World-action models (WAM) predict future states to guide robot actions, enabling learning from both action-free video and action-labeled interaction. Most inherit pretrained video generators, leaving WAM pretraining and scaling underexplored. We introduce Genie Envisioner Act 2.0 (GE-Act 2.0), a world-action model whose trainable generative and action components are all initialized from scratch on manipulation data. It combines a control-oriented autoencoder (CoAE), a single-step visual planner (SVP), and an inverse dynamics model (IDM). CoAE retains action- and instruction-relevant information under aggressive compression, while SVP produces a complete future state in one differentiable pass, so visual planning and inverse dynamics can be pretrained separately on complementary data. The components are then jointly trained with knowledge-aligned selective optimization (KASO), which reduces mismatched supervision by selecting only predicted futures judged behaviorally compatible with the recorded action. We evaluate pretrained checkpoints directly, without per-task fine-tuning, on 100 tasks across 20 manipulation skill groups with held-out scenes, backgrounds, lighting, and object instances. Scaling co-training data from 300 to 30,000 hours raises success from 17.1% to 44.1% on G1-OP and from 13.4% to 31.1% on G2-90D; despite comprising less than 2% of the co-training data, G2-90D improves by 17.7 points, suggesting cross-embodiment transfer. Gains span 19/20 and 18/20 skill groups, and skill-specific coverage strongly correlates with zero-shot out-of-distribution (OOD) success (Pearson r=0.80; Spearman rho=0.85). Under the same protocol, the model grounds object, color, shape, and position references in at least 90% of trials and follows explicit instructions even when they conflict with an already-committed behavior or a conventional scene association.