HiFi-UMI: 高忠実度UMIデータのみから展開可能な操作ポリシーを学習する
HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone
July 28, 2026
著者: Simple AI, Yuteng Wei, Jinming Ma, Jiawei Wang, Weitao Zhou, Yushen Zuo, Ke Rui, Minglei Li, Jinhao Zhang, Zhikang Pan, Xiang Wang, Haoran Jia, Huan Du, Zicheng Zeng, Jun Ma, Guiyu Qin, Di Zhang, Xiaofei Li
cs.AI
要旨
展開可能な操作ポリシーの学習は、高忠実度かつスケーラブルなデータの不足によってボトルネックとなっている。実ロボットの遠隔操作は正確だが、拡張にはコストがかかる。一方、ロボットを用いないUMIデータ収集は容易にスケールでき、現在の慣行では得られたデータを主に事前学習に用い、事後学習で少数の実ロボットによる「アンカー」を追加している。本研究では、実ロボットの割合を減らすのではなく、ロボットを用いないUMIデータの忠実度を高めることで、そのアンカーを不要にできるかどうかを問う。我々はHiFi-UMIを提案する。これは、軌跡精度、グリッパー間の相対姿勢、同期、視野を共設計したポータブルなUMIデータ生成システムである。具体的には、ヘッドマウント型オフラインステレオ慣性SLAM、再構成ではなくネイティブな相対姿勢、共有マイクロ秒GPIOトリガー、各ハンドに約200度をカバーする2台の広角カメラを備える。外部トラッキングインフラなしで、作業空間内で3mmのエンドエフェクタ精度を達成する。本データ群を用いて、ロボットを介さない事後学習(ゼロロボット事後学習)を実証する。HiFi-UMIの実演データのみで事後学習したポリシーは、実ロボットに直接展開可能であり、視覚-言語-行動ファミリーと世界-行動-モデルファミリーにわたる3つのバックボーンにおいて、同一ドメインの遠隔操作と同等の性能を示す。StarVLA-QwenPI、OpenPI-pi_0.5、LingBot-VAでの成功率差はそれぞれ-2.5、+3.1、-0.6パーセントポイントである。最も強力なポリシーは精密挿入タスクで85%の成功率を達成した。ただし、遠隔操作ベースラインは評価シーンで収集されており、HiFi-UMIの軌跡は評価シーンでは収集されていない。同一データ群からの4,000時間の事前学習により、未見の10タスクにおける行動誤差を41%削減し、StarVLA-QwenPIでは実ロボットの成功率をさらに18.1パーセントポイント向上させる。我々はHiFi-UMI-2Kをオープンソースとして公開する。これは2,000時間のマイクロ秒同期・超広視野角の実演データであり、各データは自動的に再構成され、シミュレーションリプレイによって検証済みである。これはロボット学習コミュニティのための大規模・高忠実度リソースとなる。
English
Learning deployable manipulation policies is bottlenecked by the scarcity of data that is both high-fidelity and scalable. Real-robot teleoperation is accurate but costly to scale; robot-free UMI capture scales readily, and current practice uses the resulting data mainly for pre-training, adding a small real-robot "anchor" at post-training. We ask whether raising the fidelity of robot-free UMI data, rather than shrinking the real-robot fraction, can remove that anchor. We present HiFi-UMI, a portable UMI data-production system co-designed for trajectory accuracy, inter-gripper relative pose, synchronization, and field of view: head-mounted offline stereo-inertial SLAM, native rather than reconstructed relative pose, a shared microsecond GPIO trigger, and two wide-angle cameras per hand covering ~200 degrees. It reaches 3 mm workspace-local end-effector accuracy without external tracking infrastructure. Using this corpus, we demonstrate zero-robot post-training: a policy post-trained solely on HiFi-UMI demonstrations deploys directly on a real robot and matches in-domain teleoperation across three backbones spanning the vision-language-action and world-action-model families, with success-rate differences of -2.5, +3.1, and -0.6 percentage points on StarVLA-QwenPI, OpenPI-pi_0.5, and LingBot-VA; the strongest policy reaches 85% on a precision insertion task, even though the teleoperation baseline is collected in the evaluation scene and no HiFi-UMI trajectory is. Pre-training on 4,000 hours from the same corpus lowers action error on ten unseen tasks by 41% and, on StarVLA-QwenPI, raises real-robot success by a further 18.1 percentage points. We open-source HiFi-UMI-2K, 2,000 hours of microsecond-synchronized, ultra-wide-FoV demonstrations, each automatically reconstructed and validated through simulation replay, as a large-scale, high-fidelity resource for the robot-learning community.