ChatPaper.aiChatPaper

EgoSteer: 1인칭 비디오로부터 조종 가능한 정교한 조작을 위한 풀스택 시스템

EgoSteer: A Full-Stack System Towards Steerable Dexterous Manipulation from Egocentric Videos

June 21, 2026
저자: Yifan Zhong, Zhang Chen, Tianrui Guan, Fanlian Zeng, Yuyao Ye, Tianjia He, Ka Nam Lui, Jiayi Li, Tingrui Zhang, Ruilin Yan, Xinhao Ji, Guangyu Zhao, Wenjie Lou, Jiayuan Zhang, Yuanpei Chen, Yaodong Yang
cs.AI

초록

조종 가능성은 범용 로봇 정책의 핵심 능력이지만, 대규모의 언어 정렬 및 행동 정확한 시연 데이터가 부족하여 정교한 손 시스템에서는 대부분 결여되어 있다. 이 병목 현상을 해결하기 위해, 우리는 자아 중심 인간 비디오로부터 정교한 VLA 사전 훈련을 확장하고 데이터 효율적인 실제 로봇 사후 훈련을 가능하게 하는 풀스택 시스템을 제시한다. 이 시스템은 EgoSmith(실제 환경의 자아 중심 비디오를 9.6K 시간의 고품질 사전 훈련 데이터로 큐레이션하며, 이전 최고 성능보다 9배 높은 처리량과 더 나은 정확도를 제공하는 데이터 파이프라인), 원격 조작 및 인간-루프 내 수정을 위한 통합 로봇 스택, 그리고 최적화된 인프라에서 훈련된 세계 모델 강화 VLA인 EgoSteer를 통합한다. 인간 데이터 사전 훈련은 EgoSteer에 언어 유도 조작 사전 지식을 제공하며, 이는 로봇 사후 훈련을 통해 기반화되고 DAgger 개선을 통해 향상된다. 실험적으로, EgoSteer는 40개 이상의 다양한 작업에서 자유 형식 명령을 강건하게 실행하며, 오류 복구, 정교함 및 일반화를 입증한다. 사전 훈련된 모델은 또한 복잡한 장기적 작업(박스 접기 포함)에 대해 두 가지 구현체에서 75% 이상의 성공률로 소수 샷 적응을 수행한다. 우리는 시스템, 데이터 및 모델을 https://egosteer.github.io/에서 오픈소스로 공개한다.
English
Steerability is a defining capability of generalist robot policies, yet remains largely absent in dexterous-hand systems for lack of large-scale, language-aligned, and action-accurate demonstration data. To address this bottleneck, we present a full-stack system that scales dexterous VLA pre-training from egocentric human videos and enables data-efficient real-robot post-training. It integrates EgoSmith, a data pipeline that curates in-the-wild egocentric videos into 9.6K hours of high-quality pre-training data with 9x higher throughput and better accuracy than prior SOTA; a unified robot stack for teleoperation and human-in-the-loop correction; and EgoSteer, a world-model-enhanced VLA trained on optimized infrastructure. Human-data pre-training equips EgoSteer with language-guided manipulation priors, which are grounded through robot post-training and improved by DAgger refinement. Empirically, EgoSteer robustly executes free-form instructions across 40+ diverse tasks, demonstrating failure recovery, dexterity, and generalization. The pre-trained model also few-shot adapts to complex long-horizon tasks, including box folding, on two embodiments with 75+% success. We open-source the system, data, and model at https://egosteer.github.io/.