ChatPaper.aiChatPaper

Ego2Robot:一人称視点の人間データからのスケーラブルなロボットデータ合成

Ego2Robot: Scalable Robot Data Synthesis from Egocentric Human Data

August 3, 2026
著者: Ye Wang, Pei Lin, Xiong-Hui Chen, Haoqi Yuan, Zhixuan Liang, Yiyang Huang, Anzhe Chen, Zixing Lei, Jie Zhang, Tao Zhang, Haoyang Li, Tong Zhang, Chenxi Xiao, Ziyuan Jiao, Qin Jin
cs.AI

要旨

汎化可能なロボット操作ポリシーの学習には、大規模かつ多様なデモンストレーションデータが必要である。自己中心視点の人間による操作動画は、豊富なシーンとタスクの多様性を提供し、先行研究では、そのような動画をリターゲティングしてロボット形式のデータにレンダリングすることで、小規模ながらタスクごとに効果的なポリシーを生成できることが示されている。しかし、このアプローチが視覚・言語・行動モデル(VLAモデル)に対して大規模な事前学習の利点をもたらすかどうかは未検討のままである。本稿では、アクションのリターゲティング、ロボットアームの視覚合成、多段階の品質キュレーションを通じて、自己中心視点の人間による操作動画をロボット訓練データに変換するスケーラブルなパイプラインであるEgo2Robotを提案する。Ego2Robotは、キュレーション済みデータセットと非制御環境の動画の両方をサポートし、15種類のロボット形態にわたる18,561時間のロボット訓練データを生成する。これは、現在までで最大のエゴからロボットへのデータセットである。汎化性能を評価するため、我々はRoboTwin2.0を拡張し、視覚的外観、シーンレイアウト、身体形態、タスク意味論をカバーする分離された摂動軸を導入する。実験により、Ego2Robotで合成されたデータとロボットデータによる共同事前学習が、複数の摂動タイプにわたって分布外汎化を一貫して向上させることが示され、その利点は実ロボットへの展開においても検証された。プロジェクトページ: https://www-ye.github.io/ego2robot_blog/
English
Learning generalizable robot manipulation policies requires large-scale and diverse demonstration data. Egocentric human manipulation videos offer rich scene and task diversity, and prior work has shown that retargeting and rendering such videos into robot-format data can yield effective per-task policies at small scale. However, whether this approach can provide pretraining benefits for vision-language-action models at scale remains unexplored. We present Ego2Robot, a scalable pipeline that converts egocentric human manipulation videos into robot training data through action retargeting, robot-arm visual synthesis, and multi-level quality curation. Ego2Robot supports both curated datasets and in-the-wild videos, producing 18,561 hours of robot training data spanning 15 robot morphologies, making it the largest ego-to-robot dataset to date. To evaluate generalization, we extend RoboTwin2.0 with disentangled perturbation axes covering visual appearance, scene layout, embodiment morphology, and task semantics. Experiments show that joint pretraining on Ego2Robot-synthesized and robot data consistently improves out-of-distribution generalization across multiple perturbation types, with benefits validated on real-robot deployment. Project page: https://www-ye.github.io/ego2robot_blog/