Open-AoE: 체화 학습을 위한 공개 자기중심 조작 데이터셋 및 툴체인
Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning
July 15, 2026
저자: Zishuo Li, Bowen Yang, Changtao Miao, Kai Zhu, Hao Chen, Qingze Guan, Zhengxing Wu, Wanke Zhan, Yang Sun, Zhiyi Huang, Zitong Shan, Zhenchao Jin, Jiadong Hong, Taowen Wang, Yushi Feng, You Liu, Yibo Wang, Yifan Yang, Zhaowen Zhou, Man Luo, Hao Cheng, Bo Zhang, Jianshu Li, Jiansheng Cai, Guocai Yao, Jize Zhang, Chenhao Lin, Renjing Xu, Lequan Yu, Chao Shen, Chunhua Shen, Zhe Li
cs.AI
초록
인간의 조작 행동을 담은 자기중심적 동영상은 체화된 지능을 위한 확장 가능한 감독 신호를 제공하지만, 기존 자원들은 저비용 연속 촬영, 조작 수준의 구조화된 주석, 그리고 로봇 학습을 위한 재사용 가능한 도구를 거의 결합하지 못하고 있습니다. 우리는 Open-AoE를 소개합니다. 이는 스마트폰 촬영부터 모델 학습까지 전체 파이프라인을 포괄하는 개방형, 커뮤니티 중심의 자기중심적 조작 데이터셋 및 도구체인입니다. 첫 번째 릴리스에는 400대 이상의 스마트폰을 사용한 500명 이상의 기여자들이 자연 환경에서 수집한 약 2,000시간 분량의 조작 동영상이 포함되어 있습니다. 데이터셋은 텍스트 주석, MANO 기반 손 자세, 카메라 궤적, 그리고 시간적으로 국소화된 원자적 행동을 제공합니다. Open-AoE는 또한 시간적 행동 분할, 의미 주석, 손 재구성, 카메라 궤적 재구성을 통해 원시 녹화물을 구조화된 샘플로 변환하는 데이터 처리 파이프라인을 포함합니다. 한편, 우리는 시각화, 교차 체화 리타겟팅, 모델별 데이터 변환, 그리고 VLA 정책, WAM, 월드 모델을 위한 학습 레시피를 지원하는 별도의 다운스트림 도구체인을 제공합니다. 확장 가능한 수집, 구조화된 처리, 그리고 다운스트림 적응을 통합함으로써, Open-AoE는 데이터 기여와 재사용 모두에 대한 장벽을 낮추며, 체화된 모델 학습, 인간-로봇 전이, 그리고 월드 모델링을 위한 실용적인 개방형 인프라를 제공합니다.
English
Egocentric videos of human manipulation provide scalable supervision for embodied intelligence, yet existing resources rarely combine low-cost continuous capture, manipulation-level structured annotations, and reusable tools for robot learning. We present Open-AoE, an open, community-oriented egocentric manipulation dataset and toolchain spanning the full pipeline from smartphone capture to model training. Its first release contains approximately 2,000 hours of manipulation video collected in natural environments by 500+ contributors using 400+ smartphones. The dataset provides text annotations, MANO-based hand poses, camera trajectories, and temporally localized atomic actions. Open-AoE further includes a data processing pipeline that transforms raw recordings into structured samples through temporal action segmentation, semantic annotation, hand reconstruction, and camera trajectory reconstruction. Meanwhile, we provide a separate downstream toolchain supports visualization, cross-embodiment retargeting, model-specific data conversion, and training recipes for VLA policies, WAMs, and World Models. By integrating scalable capture, structured processing, and downstream adaptation, Open-AoE reduces the barriers to both data contribution and reuse, providing practical open infrastructure for embodied model training, human-to-robot transfer, and world modeling.