ChatPaper.aiChatPaper

Open-AoE: 身体化学習のためのオープンなエゴセントリックマニピュレーションデータセットとツールチェーン

Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning

July 15, 2026
著者: Zishuo Li, Bowen Yang, Changtao Miao, Kai Zhu, Hao Chen, Qingze Guan, Zhengxing Wu, Wanke Zhan, Yang Sun, Zhiyi Huang, Zitong Shan, Zhenchao Jin, Jiadong Hong, Taowen Wang, Yushi Feng, You Liu, Yibo Wang, Yifan Yang, Zhaowen Zhou, Man Luo, Hao Cheng, Bo Zhang, Jianshu Li, Jiansheng Cai, Guocai Yao, Jize Zhang, Chenhao Lin, Renjing Xu, Lequan Yu, Chao Shen, Chunhua Shen, Zhe Li
cs.AI

要旨

人間の操作を記録した自己中心視点の映像は、身体化知能に対するスケーラブルな教師信号を提供するが、既存のリソースは、低コストでの継続的な収集、操作レベルの構造化アノテーション、ロボット学習のための再利用可能なツールを兼ね備えることはほとんどない。本稿では、スマートフォンによる撮影からモデル訓練に至る全パイプラインをカバーする、オープンでコミュニティ指向の自己中心視点操作データセットとツールチェーンであるOpen-AoEを提案する。初回リリースでは、500名以上の投稿者が400台以上のスマートフォンを用いて自然環境で収集した、約2,000時間の操作映像を含む。データセットには、テキストアノテーション、MANOベースの手姿勢、カメラ軌跡、時間的に局所化された原子動作が提供される。さらにOpen-AoEは、時間的行動セグメンテーション、意味アノテーション、手の再構築、カメラ軌跡再構築により、生の記録を構造化サンプルに変換するデータ処理パイプラインを備える。同時に、別途提供される下流タスク用ツールチェーンは、可視化、異なる身体へのリターゲティング、モデル固有のデータ変換、そしてVLAポリシー、WAM、世界モデル向けの学習レシピをサポートする。スケーラブルな収集、構造化処理、下流タスクへの適応を統合することにより、Open-AoEはデータ提供と再利用の両方における障壁を低減し、身体化モデル訓練、人間からロボットへの転移、世界モデリングのための実用的なオープンインフラを提供する。
English
Egocentric videos of human manipulation provide scalable supervision for embodied intelligence, yet existing resources rarely combine low-cost continuous capture, manipulation-level structured annotations, and reusable tools for robot learning. We present Open-AoE, an open, community-oriented egocentric manipulation dataset and toolchain spanning the full pipeline from smartphone capture to model training. Its first release contains approximately 2,000 hours of manipulation video collected in natural environments by 500+ contributors using 400+ smartphones. The dataset provides text annotations, MANO-based hand poses, camera trajectories, and temporally localized atomic actions. Open-AoE further includes a data processing pipeline that transforms raw recordings into structured samples through temporal action segmentation, semantic annotation, hand reconstruction, and camera trajectory reconstruction. Meanwhile, we provide a separate downstream toolchain supports visualization, cross-embodiment retargeting, model-specific data conversion, and training recipes for VLA policies, WAMs, and World Models. By integrating scalable capture, structured processing, and downstream adaptation, Open-AoE reduces the barriers to both data contribution and reuse, providing practical open infrastructure for embodied model training, human-to-robot transfer, and world modeling.