EgoSteer: 一人称視点動画からの誘導可能な巧緻操作を実現するフルスタックシステム
EgoSteer: A Full-Stack System Towards Steerable Dexterous Manipulation from Egocentric Videos
June 21, 2026
著者: Yifan Zhong, Zhang Chen, Tianrui Guan, Fanlian Zeng, Yuyao Ye, Tianjia He, Ka Nam Lui, Jiayi Li, Tingrui Zhang, Ruilin Yan, Xinhao Ji, Guangyu Zhao, Wenjie Lou, Jiayuan Zhang, Yuanpei Chen, Yaodong Yang
cs.AI
要旨
ステアラビリティは汎用ロボットポリシーの定義的能力であるが、大規模で言語に整合し、動作精度の高いデモンストレーションデータが不足しているため、器用なハンドシステムではほとんど実現されていない。このボトルネックを解決するために、我々は、人間の主体視点映像から器用なVLA(Vision-Language-Action)の事前学習をスケールさせ、データ効率の良い実ロボットの事後学習を可能にするフルスタックシステムを提案する。このシステムは、野外の主体視点映像を厳選して9,600時間の高品質な事前学習データに変換するデータパイプラインEgoSmith(先行SOTA比で9倍のスループットとより高い精度を達成)、遠隔操作と人間参加型補正のための統合ロボットスタック、最適化されたインフラ上で訓練された世界モデル強化VLAであるEgoSteerを統合している。人間データによる事前学習は、EgoSteerに言語誘導操作の事前知識を与え、これはロボット事後学習で基盤化され、DAggerによる洗練によって改善される。経験的に、EgoSteerは40種類以上の多様なタスクにわたって自由形式の指示をロバストに実行し、障害からの回復、器用さ、一般化を示している。この事前学習モデルは、箱折りを含む複雑な長期的タスクにも、2つの身体で75%以上の成功率で少数ショット適応する。我々はシステム、データ、モデルを https://egosteer.github.io/ でオープンソース化している。
English
Steerability is a defining capability of generalist robot policies, yet remains largely absent in dexterous-hand systems for lack of large-scale, language-aligned, and action-accurate demonstration data. To address this bottleneck, we present a full-stack system that scales dexterous VLA pre-training from egocentric human videos and enables data-efficient real-robot post-training. It integrates EgoSmith, a data pipeline that curates in-the-wild egocentric videos into 9.6K hours of high-quality pre-training data with 9x higher throughput and better accuracy than prior SOTA; a unified robot stack for teleoperation and human-in-the-loop correction; and EgoSteer, a world-model-enhanced VLA trained on optimized infrastructure. Human-data pre-training equips EgoSteer with language-guided manipulation priors, which are grounded through robot post-training and improved by DAgger refinement. Empirically, EgoSteer robustly executes free-form instructions across 40+ diverse tasks, demonstrating failure recovery, dexterity, and generalization. The pre-trained model also few-shot adapts to complex long-horizon tasks, including box folding, on two embodiments with 75+% success. We open-source the system, data, and model at https://egosteer.github.io/.