GE-Act 2.0: 로봇 조작을 위한 월드-액션 모델의 사전 학습 및 스케일링
GE-Act 2.0: Pretraining and Scaling a World-Action Model for Robotic Manipulation
September 4, 2026
저자: AgiBot Research Team, Renhang Liu, Wenzhi Zhao, Zhuo Yang, Liliang Chen, Pengfei Zhou, Shengcong Chen, Guanghui Ren, Youlun Peng, Rongjun Jin, Nan Wang, Sukai Wang, Xindong He, Jinyuan Feng, Ziyu Xiong, Linqing Zhong, Yifei Wei, Feng Han, Long Zhang, Da Huang, Nanshu Zhao, Chenghao Yin, Mo Wu, Zhaodong Yan, Kongtao Hu, Yuxiang Yan, Aogelijiang Niyazi, Yu Fang, Jia Zeng, Lizhu Meng, Daizhen Lv, Haoyu Cao, Zhiwen Hou, Lianjin Ye, Yuehan Niu, Zhikai Cai, Xuan Hu, Hui Min, Xiongfeng Cai, Yue Liao, Jing Wu, Soujanya Poria, Ye Li, Sanping Zhou, Maoqing Yao
cs.AI
초록
월드-액션 모델(WAM)은 미래 상태를 예측하여 로봇 행동을 안내하며, 행동이 없는 비디오와 행동 레이블이 있는 상호작용 모두로부터 학습할 수 있게 한다. 대부분은 사전학습된 비디오 생성기를 계승하므로, WAM의 사전학습과 스케일링은 충분히 탐구되지 않은 채로 남아 있다. 우리는 학습 가능한 생성 및 행동 구성요소가 모두 매니퓰레이션 데이터에서 처음부터 초기화되는 월드-액션 모델인 Genie Envisioner Act 2.0(GE-Act 2.0)을 소개한다. 이는 제어 지향 오토인코더(CoAE), 단일 스텝 시각 플래너(SVP), 역동역학 모델(IDM)을 결합한다. CoAE는 강한 압축 하에서도 행동 및 지시 관련 정보를 보존하는 반면, SVP는 미분 가능한 단일 패스로 완전한 미래 상태를 생성하므로, 시각 플래닝과 역동역학은 상보적 데이터에서 별도로 사전학습될 수 있다. 이후 구성요소들은 지식 정렬 선택적 최적화(KASO)를 통해 공동 학습되는데, 이는 기록된 행동과 행동적으로 양립한다고 판단된 예측 미래만을 선택함으로써 감독 불일치를 줄인다. 우리는 홀드아웃 장면, 배경, 조명, 객체 인스턴스를 사용한 20개 매니퓰레이션 기술 그룹 전반의 100개 작업에서 작업별 미세 조정 없이 사전학습된 체크포인트를 직접 평가한다. 공동 학습 데이터를 300시간에서 30,000시간으로 스케일링하면 G1-OP에서 성공률이 17.1%에서 44.1%로, G2-90D에서 13.4%에서 31.1%로 향상된다. G2-90D는 공동 학습 데이터의 2% 미만을 차지함에도 17.7점 향상되어 교차 엠바디먼트 전이를 시사한다. 성능 향상은 19/20 및 18/20 기술 그룹에 걸쳐 나타났으며, 기술별 커버리지는 제로샷 분포 외(OOD) 성공률과 강한 상관관계를 보였다(Pearson r=0.80; Spearman rho=0.85). 동일한 프로토콜에서 이 모델은 최소 90%의 시행에서 객체, 색상, 모양, 위치 참조를 그라운딩하며, 이미 확정된 행동이나 관습적인 장면 연상과 충돌하는 경우에도 명시적 지시를 따른다.
English
World-action models (WAM) predict future states to guide robot actions, enabling learning from both action-free video and action-labeled interaction. Most inherit pretrained video generators, leaving WAM pretraining and scaling underexplored. We introduce Genie Envisioner Act 2.0 (GE-Act 2.0), a world-action model whose trainable generative and action components are all initialized from scratch on manipulation data. It combines a control-oriented autoencoder (CoAE), a single-step visual planner (SVP), and an inverse dynamics model (IDM). CoAE retains action- and instruction-relevant information under aggressive compression, while SVP produces a complete future state in one differentiable pass, so visual planning and inverse dynamics can be pretrained separately on complementary data. The components are then jointly trained with knowledge-aligned selective optimization (KASO), which reduces mismatched supervision by selecting only predicted futures judged behaviorally compatible with the recorded action. We evaluate pretrained checkpoints directly, without per-task fine-tuning, on 100 tasks across 20 manipulation skill groups with held-out scenes, backgrounds, lighting, and object instances. Scaling co-training data from 300 to 30,000 hours raises success from 17.1% to 44.1% on G1-OP and from 13.4% to 31.1% on G2-90D; despite comprising less than 2% of the co-training data, G2-90D improves by 17.7 points, suggesting cross-embodiment transfer. Gains span 19/20 and 18/20 skill groups, and skill-specific coverage strongly correlates with zero-shot out-of-distribution (OOD) success (Pearson r=0.80; Spearman rho=0.85). Under the same protocol, the model grounds object, color, shape, and position references in at least 90% of trials and follows explicit instructions even when they conflict with an already-committed behavior or a conventional scene association.