GE-Act 2.0: ロボットマニピュレーションのためのワールド・アクション・モデルの事前学習とスケーリング
GE-Act 2.0: Pretraining and Scaling a World-Action Model for Robotic Manipulation
September 4, 2026
著者: AgiBot Research Team, Renhang Liu, Wenzhi Zhao, Zhuo Yang, Liliang Chen, Pengfei Zhou, Shengcong Chen, Guanghui Ren, Youlun Peng, Rongjun Jin, Nan Wang, Sukai Wang, Xindong He, Jinyuan Feng, Ziyu Xiong, Linqing Zhong, Yifei Wei, Feng Han, Long Zhang, Da Huang, Nanshu Zhao, Chenghao Yin, Mo Wu, Zhaodong Yan, Kongtao Hu, Yuxiang Yan, Aogelijiang Niyazi, Yu Fang, Jia Zeng, Lizhu Meng, Daizhen Lv, Haoyu Cao, Zhiwen Hou, Lianjin Ye, Yuehan Niu, Zhikai Cai, Xuan Hu, Hui Min, Xiongfeng Cai, Yue Liao, Jing Wu, Soujanya Poria, Ye Li, Sanping Zhou, Maoqing Yao
cs.AI
要旨
ワールドアクションモデル(WAM)は将来状態を予測してロボット行動を導き、行動ラベルなし動画と行動ラベル付きインタラクションの両方から学習することを可能にする。そのほとんどは事前学習済み動画生成器を継承しており、WAMの事前学習とスケーリングは十分に検討されていないままである。我々はGenie Envisioner Act 2.0(GE-Act 2.0)を導入する。これは、学習可能な生成コンポーネントと行動コンポーネントのすべてをマニピュレーションデータでスクラッチから初期化するワールドアクションモデルである。これは制御指向オートエンコーダ(CoAE)、単一ステップ視覚プランナー(SVP)、逆動力学モデル(IDM)を組み合わせる。CoAEは強力な圧縮の下でも行動および指示に関連する情報を保持し、SVPは1回の微分可能なパスで完全な将来状態を生成する。そのため、視覚プランニングと逆動力学は相補的なデータで別々に事前学習できる。次に、各コンポーネントは知識整合型選択的最適化(KASO)を用いて統合訓練される。KASOは、記録された行動と行動的に整合すると判断された予測将来のみを選択することで、不一致な教師信号を低減する。我々は、タスクごとの微調整なしに、事前学習済みチェックポイントを直接評価する。評価は、ホールドアウトされたシーン、背景、照明、物体インスタンスを用いて、20のマニピュレーションスキルグループにわたる100タスクで行う。共訓練データを300時間から30,000時間にスケールすると、G1-OPでは成功率が17.1%から44.1%に、G2-90Dでは13.4%から31.1%に上昇する。G2-90Dは共訓練データの2%未満しか占めないにもかかわらず17.7ポイント改善しており、クロスエンボディメント転移を示唆する。改善は19/20および18/20のスキルグループに及び、スキル固有のカバレッジはゼロショット分布外(OOD)成功率と強く相関する(Pearson r=0.80、Spearman rho=0.85)。同一のプロトコルの下で、モデルは少なくとも90%の試行で物体、色、形状、位置の参照をグラウンディングし、明示的な指示がすでに確定した行動や慣例的なシーン連想と矛盾する場合でもそれに従う。
English
World-action models (WAM) predict future states to guide robot actions, enabling learning from both action-free video and action-labeled interaction. Most inherit pretrained video generators, leaving WAM pretraining and scaling underexplored. We introduce Genie Envisioner Act 2.0 (GE-Act 2.0), a world-action model whose trainable generative and action components are all initialized from scratch on manipulation data. It combines a control-oriented autoencoder (CoAE), a single-step visual planner (SVP), and an inverse dynamics model (IDM). CoAE retains action- and instruction-relevant information under aggressive compression, while SVP produces a complete future state in one differentiable pass, so visual planning and inverse dynamics can be pretrained separately on complementary data. The components are then jointly trained with knowledge-aligned selective optimization (KASO), which reduces mismatched supervision by selecting only predicted futures judged behaviorally compatible with the recorded action. We evaluate pretrained checkpoints directly, without per-task fine-tuning, on 100 tasks across 20 manipulation skill groups with held-out scenes, backgrounds, lighting, and object instances. Scaling co-training data from 300 to 30,000 hours raises success from 17.1% to 44.1% on G1-OP and from 13.4% to 31.1% on G2-90D; despite comprising less than 2% of the co-training data, G2-90D improves by 17.7 points, suggesting cross-embodiment transfer. Gains span 19/20 and 18/20 skill groups, and skill-specific coverage strongly correlates with zero-shot out-of-distribution (OOD) success (Pearson r=0.80; Spearman rho=0.85). Under the same protocol, the model grounds object, color, shape, and position references in at least 90% of trials and follows explicit instructions even when they conflict with an already-committed behavior or a conventional scene association.