ChatPaper.aiChatPaper

HiFi-UMI:僅從高保真UMI資料學習可部署的操作策略

HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone

July 28, 2026
作者: Simple AI, Yuteng Wei, Jinming Ma, Jiawei Wang, Weitao Zhou, Yushen Zuo, Ke Rui, Minglei Li, Jinhao Zhang, Zhikang Pan, Xiang Wang, Haoran Jia, Huan Du, Zicheng Zeng, Jun Ma, Guiyu Qin, Di Zhang, Xiaofei Li
cs.AI

摘要

學習可部署的操作策略面臨瓶頸,原因在於同時具備高保真度與可擴展性的資料極為稀缺。真實機器人遙控操作精準但擴展成本高昂;無機器人UMI資料收集則易於擴展,現行實務主要將所得資料用於預訓練,並在後訓練階段加入少量真實機器人資料作為「錨點」。我們探討:提升無機器人UMI資料的保真度(而非縮減真實機器人資料比例),是否能移除該錨點。我們提出HiFi-UMI,這是一套可攜式UMI資料生產系統,針對軌跡精度、夾爪間相對位姿、同步性與視野共同設計:頭戴式離線立體慣性SLAM、原生而非重建的相對位姿、共享微秒級GPIO觸發器,以及每隻手配備兩顆涵蓋約200度的廣角攝影機。該系統在無外部追蹤基礎設施下,可達到工作空間內末端執行器3公釐的局部精度。利用此資料庫,我們展示零機器人後訓練:僅以HiFi-UMI示範資料進行後訓練的策略,可直接部署於真實機器人,並在三個涵蓋視覺-語言-動作與世界-動作模型系列的骨幹架構上,與領域內遙控操作表現相當——在StarVLA-QwenPI、OpenPI-pi_0.5與LingBot-VA上的成功率差異分別為-2.5、+3.1與-0.6個百分點;最強的策略在一項精密插入任務上達到85%的成功率,儘管遙控操作基準線是在評估場景中收集的,而無任何HiFi-UMI軌跡來自該場景。使用同一資料庫中4,000小時的資料進行預訓練,在十項未見過任務上的動作誤差降低了41%,並在StarVLA-QwenPI上進一步將真實機器人成功率提升18.1個百分點。我們開源HiFi-UMI-2K,包含2,000小時微秒同步、超廣視野的示範資料,每段資料皆自動重建並透過模擬重演驗證,為機器人學習社群提供大規模、高保真度的資源。
English
Learning deployable manipulation policies is bottlenecked by the scarcity of data that is both high-fidelity and scalable. Real-robot teleoperation is accurate but costly to scale; robot-free UMI capture scales readily, and current practice uses the resulting data mainly for pre-training, adding a small real-robot "anchor" at post-training. We ask whether raising the fidelity of robot-free UMI data, rather than shrinking the real-robot fraction, can remove that anchor. We present HiFi-UMI, a portable UMI data-production system co-designed for trajectory accuracy, inter-gripper relative pose, synchronization, and field of view: head-mounted offline stereo-inertial SLAM, native rather than reconstructed relative pose, a shared microsecond GPIO trigger, and two wide-angle cameras per hand covering ~200 degrees. It reaches 3 mm workspace-local end-effector accuracy without external tracking infrastructure. Using this corpus, we demonstrate zero-robot post-training: a policy post-trained solely on HiFi-UMI demonstrations deploys directly on a real robot and matches in-domain teleoperation across three backbones spanning the vision-language-action and world-action-model families, with success-rate differences of -2.5, +3.1, and -0.6 percentage points on StarVLA-QwenPI, OpenPI-pi_0.5, and LingBot-VA; the strongest policy reaches 85% on a precision insertion task, even though the teleoperation baseline is collected in the evaluation scene and no HiFi-UMI trajectory is. Pre-training on 4,000 hours from the same corpus lowers action error on ten unseen tasks by 41% and, on StarVLA-QwenPI, raises real-robot success by a further 18.1 percentage points. We open-source HiFi-UMI-2K, 2,000 hours of microsecond-synchronized, ultra-wide-FoV demonstrations, each automatically reconstructed and validated through simulation replay, as a large-scale, high-fidelity resource for the robot-learning community.