ChatPaper.aiChatPaper

Open-AoE:面向具身学习的开放第一人称操作数据集与工具链

Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning

July 15, 2026
作者: Zishuo Li, Bowen Yang, Changtao Miao, Kai Zhu, Hao Chen, Qingze Guan, Zhengxing Wu, Wanke Zhan, Yang Sun, Zhiyi Huang, Zitong Shan, Zhenchao Jin, Jiadong Hong, Taowen Wang, Yushi Feng, You Liu, Yibo Wang, Yifan Yang, Zhaowen Zhou, Man Luo, Hao Cheng, Bo Zhang, Jianshu Li, Jiansheng Cai, Guocai Yao, Jize Zhang, Chenhao Lin, Renjing Xu, Lequan Yu, Chao Shen, Chunhua Shen, Zhe Li
cs.AI

摘要

以人类操作为中心的第一视角视频为具身智能提供了可扩展的监督信号,然而现有资源很少能同时兼顾低成本连续采集、操作级别的结构化标注以及可复用的机器人学习工具。我们提出Open-AoE——一个开放、面向社区的第一视角操作数据集与工具链,涵盖从智能手机拍摄到模型训练的完整流程。其首个版本包含约2000小时的操作视频,由500多名贡献者使用400余部智能手机在自然环境中采集。该数据集提供文本标注、基于MANO的手部姿态、相机轨迹以及时间定位的原子动作。Open-AoE还包含一个数据处理流水线,通过时序动作分割、语义标注、手部重建和相机轨迹重建,将原始录制内容转化为结构化样本。同时,我们提供了一个独立的下游工具链,支持可视化、跨体态重定向、模型特定数据转换,以及针对视觉语言-动作策略(VLA)、世界行动模型(WAMs)和世界模型的训练方案。通过整合可扩展采集、结构化处理与下游适配,Open-AoE降低了数据贡献与复用的门槛,为具身模型训练、人机转移和世界建模提供了实用的开放基础设施。
English
Egocentric videos of human manipulation provide scalable supervision for embodied intelligence, yet existing resources rarely combine low-cost continuous capture, manipulation-level structured annotations, and reusable tools for robot learning. We present Open-AoE, an open, community-oriented egocentric manipulation dataset and toolchain spanning the full pipeline from smartphone capture to model training. Its first release contains approximately 2,000 hours of manipulation video collected in natural environments by 500+ contributors using 400+ smartphones. The dataset provides text annotations, MANO-based hand poses, camera trajectories, and temporally localized atomic actions. Open-AoE further includes a data processing pipeline that transforms raw recordings into structured samples through temporal action segmentation, semantic annotation, hand reconstruction, and camera trajectory reconstruction. Meanwhile, we provide a separate downstream toolchain supports visualization, cross-embodiment retargeting, model-specific data conversion, and training recipes for VLA policies, WAMs, and World Models. By integrating scalable capture, structured processing, and downstream adaptation, Open-AoE reduces the barriers to both data contribution and reuse, providing practical open infrastructure for embodied model training, human-to-robot transfer, and world modeling.