ChatPaper.aiChatPaper

Open-AoE:一個用於具身學習的開放式自我中心操作數據集與工具鏈

Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning

July 15, 2026
作者: Zishuo Li, Bowen Yang, Changtao Miao, Kai Zhu, Hao Chen, Qingze Guan, Zhengxing Wu, Wanke Zhan, Yang Sun, Zhiyi Huang, Zitong Shan, Zhenchao Jin, Jiadong Hong, Taowen Wang, Yushi Feng, You Liu, Yibo Wang, Yifan Yang, Zhaowen Zhou, Man Luo, Hao Cheng, Bo Zhang, Jianshu Li, Jiansheng Cai, Guocai Yao, Jize Zhang, Chenhao Lin, Renjing Xu, Lequan Yu, Chao Shen, Chunhua Shen, Zhe Li
cs.AI

摘要

以自我為中心的人類操作影片為具身智能提供了可擴展的監督訊號,然而現有資源鮮少同時具備低成本連續捕捉、操作層級的結構化標註以及可用於機器人學習的即用工具。我們提出Open-AoE,這是一個開放且以社群為導向的自我中心操作資料集與工具鏈,涵蓋從智慧型手機拍攝到模型訓練的完整流程。其首個釋出版本包含由500多位貢獻者使用400多支智慧型手機在自然環境中收集的約2,000小時操作影片。資料集提供了文字標註、基於MANO的手部姿態、相機軌跡以及時序定位的原子動作。Open-AoE還包含一套資料處理流程,透過時序動作分割、語意標註、手部重建與相機軌跡重建,將原始錄影轉換為結構化樣本。同時,我們提供了一套獨立的下游工具鏈,支援視覺化、跨具象重定向、模型特定資料轉換,以及針對VLA策略、WAMs與世界模型的訓練配方。透過整合可擴展的捕捉、結構化處理與下游適應,Open-AoE降低了資料貢獻與重複使用的門檻,為具身模型訓練、人機轉移與世界建模提供了實用的開放基礎設施。
English
Egocentric videos of human manipulation provide scalable supervision for embodied intelligence, yet existing resources rarely combine low-cost continuous capture, manipulation-level structured annotations, and reusable tools for robot learning. We present Open-AoE, an open, community-oriented egocentric manipulation dataset and toolchain spanning the full pipeline from smartphone capture to model training. Its first release contains approximately 2,000 hours of manipulation video collected in natural environments by 500+ contributors using 400+ smartphones. The dataset provides text annotations, MANO-based hand poses, camera trajectories, and temporally localized atomic actions. Open-AoE further includes a data processing pipeline that transforms raw recordings into structured samples through temporal action segmentation, semantic annotation, hand reconstruction, and camera trajectory reconstruction. Meanwhile, we provide a separate downstream toolchain supports visualization, cross-embodiment retargeting, model-specific data conversion, and training recipes for VLA policies, WAMs, and World Models. By integrating scalable capture, structured processing, and downstream adaptation, Open-AoE reduces the barriers to both data contribution and reuse, providing practical open infrastructure for embodied model training, human-to-robot transfer, and world modeling.