HiFi-UMI:仅从高保真UMI数据中学习可部署的操作策略
HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone
July 28, 2026
作者: Simple AI, Yuteng Wei, Jinming Ma, Jiawei Wang, Weitao Zhou, Yushen Zuo, Ke Rui, Minglei Li, Jinhao Zhang, Zhikang Pan, Xiang Wang, Haoran Jia, Huan Du, Zicheng Zeng, Jun Ma, Guiyu Qin, Di Zhang, Xiaofei Li
cs.AI
摘要
学习可部署的操作策略面临瓶颈,即缺乏同时具备高保真度和可扩展性的数据。真实机器人遥操作精确但扩展成本高昂;无机器人UMI数据采集易于扩展,当前实践中主要将产生的数据用于预训练,并在后训练阶段添加少量真实机器人数据作为“锚点”。我们探究:是否可以通过提升无机器人UMI数据的保真度(而非减少真实机器人数据的比例)来消除这一锚点。我们提出了HiFi-UMI,这是一个便携式UMI数据生产系统,其在轨迹精度、夹爪间相对位姿、同步性和视场方面进行了协同设计:头戴式离线双目惯性SLAM、原生而非重建的相对位姿、共享微秒级GPIO触发器,以及每只手配备两个覆盖约200度的广角摄像头。该系统无需外部追踪基础设施,即可实现3毫米级别的工作空间局部末端执行器精度。利用这一数据语料库,我们展示了零机器人后训练:仅基于HiFi-UMI演示数据进行后训练的策略可直接部署到真实机器人上,并在涵盖视觉-语言-动作和世界-动作模型家族的三个主干网络上,与领域内遥操作的性能相当;在StarVLA-QwenPI、OpenPI-pi_0.5和LingBot-VA上,成功率差异分别为-2.5、+3.1和-0.6个百分点。最强的策略在精密插入任务上达到85%的成功率,尽管遥操作基线在评估场景中采集,而HiFi-UMI轨迹从未在该场景中出现。在同一语料库上进行4000小时的预训练,将在十个未见任务上的动作误差降低41%,并在StarVLA-QwenPI上进一步提升真实机器人成功率18.1个百分点。我们开源了HiFi-UMI-2K,包含2000小时微秒级同步、超宽视场的演示数据,每条数据均通过仿真回放自动重建和验证,为机器人学习社区提供了一个大规模、高保真度的资源。
English
Learning deployable manipulation policies is bottlenecked by the scarcity of data that is both high-fidelity and scalable. Real-robot teleoperation is accurate but costly to scale; robot-free UMI capture scales readily, and current practice uses the resulting data mainly for pre-training, adding a small real-robot "anchor" at post-training. We ask whether raising the fidelity of robot-free UMI data, rather than shrinking the real-robot fraction, can remove that anchor. We present HiFi-UMI, a portable UMI data-production system co-designed for trajectory accuracy, inter-gripper relative pose, synchronization, and field of view: head-mounted offline stereo-inertial SLAM, native rather than reconstructed relative pose, a shared microsecond GPIO trigger, and two wide-angle cameras per hand covering ~200 degrees. It reaches 3 mm workspace-local end-effector accuracy without external tracking infrastructure. Using this corpus, we demonstrate zero-robot post-training: a policy post-trained solely on HiFi-UMI demonstrations deploys directly on a real robot and matches in-domain teleoperation across three backbones spanning the vision-language-action and world-action-model families, with success-rate differences of -2.5, +3.1, and -0.6 percentage points on StarVLA-QwenPI, OpenPI-pi_0.5, and LingBot-VA; the strongest policy reaches 85% on a precision insertion task, even though the teleoperation baseline is collected in the evaluation scene and no HiFi-UMI trajectory is. Pre-training on 4,000 hours from the same corpus lowers action error on ten unseen tasks by 41% and, on StarVLA-QwenPI, raises real-robot success by a further 18.1 percentage points. We open-source HiFi-UMI-2K, 2,000 hours of microsecond-synchronized, ultra-wide-FoV demonstrations, each automatically reconstructed and validated through simulation replay, as a large-scale, high-fidelity resource for the robot-learning community.