ChatPaper.aiChatPaper

HiFi-UMI: 고충실도 UMI 데이터만으로 배포 가능한 조작 정책 학습

HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone

July 28, 2026
저자: Simple AI, Yuteng Wei, Jinming Ma, Jiawei Wang, Weitao Zhou, Yushen Zuo, Ke Rui, Minglei Li, Jinhao Zhang, Zhikang Pan, Xiang Wang, Haoran Jia, Huan Du, Zicheng Zeng, Jun Ma, Guiyu Qin, Di Zhang, Xiaofei Li
cs.AI

초록

배치 가능한 조작 정책을 학습하는 데 있어 고충실도와 확장성을 모두 갖춘 데이터의 부족이 병목 현상이 되고 있습니다. 실제 로봇 원격 조작은 정확하지만 확장 비용이 높은 반면, 로봇 없는 UMI 캡처는 쉽게 확장 가능하며, 현재 관행은 그 결과 데이터를 주로 사전 학습에 사용하고 사후 학습에 소량의 실제 로봇 '앵커'를 추가하는 것입니다. 우리는 실제 로봇 비율을 줄이는 대신 로봇 없는 UMI 데이터의 충실도를 높이는 것이 그 앵커를 제거할 수 있는지 묻습니다. 우리는 궤적 정확도, 그리퍼 간 상대 자세, 동기화, 시야를 위해 공동 설계된 휴대용 UMI 데이터 생성 시스템인 HiFi-UMI를 제시합니다: 머리 장착형 오프라인 스테레오-관성 SLAM, 재구성된 것이 아닌 네이티브 상대 자세, 공유 마이크로초 GPIO 트리거, 각 손당 약 200도를 커버하는 두 개의 광각 카메라를 사용합니다. 외부 추적 인프라 없이 작업 공간 내 엔드 이펙터 정확도 3mm를 달성합니다. 이 코퍼스를 사용하여 우리는 제로 로봇 사후 학습을 입증합니다: HiFi-UMI 데모로만 사후 학습된 정책이 실제 로봇에 직접 배포되어 비전-언어-행동 및 세계-행동-모델 계열에 걸친 세 가지 백본에서 인도메인 원격 조작과 일치하는 성능을 보이며, StarVLA-QwenPI, OpenPI-pi_0.5, LingBot-VA에서 성공률 차이가 각각 -2.5, +3.1, -0.6 퍼센트 포인트입니다. 가장 강력한 정책은 정밀 삽입 작업에서 85%의 성공률을 달성하는데, 원격 조작 기준선은 평가 장면에서 수집되었고 HiFi-UMI 궤적은 하나도 포함되지 않았습니다. 동일한 코퍼스의 4,000시간 데이터로 사전 학습하면 보지 못한 열 가지 작업에서 행동 오류가 41% 감소하고, StarVLA-QwenPI에서는 실제 로봇 성공률이 추가로 18.1 퍼센트 포인트 향상됩니다. 우리는 로봇 학습 커뮤니티를 위한 대규모 고충실도 리소스로서 마이크로초 동기화된 초광각 시야 데모 2,000시간 분량의 HiFi-UMI-2K를 오픈소스로 공개합니다. 각 데모는 자동으로 재구성되고 시뮬레이션 재생을 통해 검증됩니다.
English
Learning deployable manipulation policies is bottlenecked by the scarcity of data that is both high-fidelity and scalable. Real-robot teleoperation is accurate but costly to scale; robot-free UMI capture scales readily, and current practice uses the resulting data mainly for pre-training, adding a small real-robot "anchor" at post-training. We ask whether raising the fidelity of robot-free UMI data, rather than shrinking the real-robot fraction, can remove that anchor. We present HiFi-UMI, a portable UMI data-production system co-designed for trajectory accuracy, inter-gripper relative pose, synchronization, and field of view: head-mounted offline stereo-inertial SLAM, native rather than reconstructed relative pose, a shared microsecond GPIO trigger, and two wide-angle cameras per hand covering ~200 degrees. It reaches 3 mm workspace-local end-effector accuracy without external tracking infrastructure. Using this corpus, we demonstrate zero-robot post-training: a policy post-trained solely on HiFi-UMI demonstrations deploys directly on a real robot and matches in-domain teleoperation across three backbones spanning the vision-language-action and world-action-model families, with success-rate differences of -2.5, +3.1, and -0.6 percentage points on StarVLA-QwenPI, OpenPI-pi_0.5, and LingBot-VA; the strongest policy reaches 85% on a precision insertion task, even though the teleoperation baseline is collected in the evaluation scene and no HiFi-UMI trajectory is. Pre-training on 4,000 hours from the same corpus lowers action error on ten unseen tasks by 41% and, on StarVLA-QwenPI, raises real-robot success by a further 18.1 percentage points. We open-source HiFi-UMI-2K, 2,000 hours of microsecond-synchronized, ultra-wide-FoV demonstrations, each automatically reconstructed and validated through simulation replay, as a large-scale, high-fidelity resource for the robot-learning community.