CosmoH2G:複雑な空間運動を伴う物体操作のためのハンドからグリッパへの転移データセットとベースライン手法
CosmoH2G: A Hand-to-Gripper Transfer Dataset and Baseline Method for Object Manipulation with Complex Spatial Movements
September 7, 2026
著者: Hongxiang Zhao, Mutian Xu, Zeyu Jin, Yiming Hao, Shuguang Cui, Xiaoguang Han
cs.AI
要旨
人間の手による実演をロボットグリッパへ転移することは、近年、ロボット学習における費用対効果の高い解決策として登場している。しかし、既存手法は主に単純な平面タスクに限定されており、ロボットマニピュレーションに不可欠な複雑な空間動作(例:回転や反転を伴う複雑な軌道)を扱えない。このギャップに動機づけられ、我々は細粒度の手姿勢動作に導かれる暗黙的かつデータ駆動型のアプローチを採用する。そのために、我々は手とグリッパのペアデモンストレーションを収集するスケーラブルな取得パイプラインを導入する。これは、動作の複雑さを優先し、シームレスな動作模倣のためにハンドヘルドグリッパを活用する厳密なプロトコルに従う。これにより、1,254個の固有物体にわたる6,189エピソードからなる大規模ペアデータセットが得られ、既存ベンチマークより著しく高い空間的複雑さを示す。しかし、そのような複雑な写像を学習することは依然として困難である。我々は、グリッパ姿勢系列全体を単純にエンドツーエンド生成するだけでは不十分であることを観察する。なぜなら、わずかな軌道偏差が複雑なダイナミクスの下で急速に累積するからである。これに対処するため、我々は二段階フレームワークを提案する。第I段階では写像目的を簡素化するために疎なグリッパキーフレーム(初期および終端)を予測し、第II段階ではこれらのキーフレームを条件として完全な連続行動系列を生成する。さらに、累積ドリフトを緩和するため、グリッパの向きは学習させつつ、その並進を把持ヒューリスティックと運動学的整合性に基づいて事後最適化する。シミュレーションと実機ロボット実験の両方において、我々のフレームワークは複雑な空間マニピュレーションの安定かつ精密な手からグリッパへの転移を可能にし、従来のベースラインを大幅に上回る。プロジェクトページ: https://cosmoh2g.github.io.
English
Transferring human hand demonstrations to robotic grippers has recently emerged as a cost-effective solution for robot learning. However, existing methods are largely confined to simple, planar tasks and fail to handle complex spatial movements (e.g., intricate trajectories involving rotations or flips) that are essential for robot manipulation. Motivated by this gap, we adopt an implicit, data-driven approach guided by fine-grained hand-pose motions. To this end, we introduce a scalable acquisition pipeline to collect hand-gripper paired demonstrations, governed by a rigorous protocol that prioritizes motion complexity and leverages a handheld gripper for seamless action mimicry. This yields a large-scale paired dataset comprising 6,189 episodes across 1,254 unique objects, exhibiting significantly higher spatial complexity than existing benchmarks. However, learning such complex mappings remains challenging. We observe that naive end-to-end generation of full gripper pose sequences is insufficient, as minor trajectory deviations compound rapidly under intricate dynamics. To address this, we propose a two-stage framework: Stage I predicts sparse gripper keyframes (initial and terminal) to simplify the mapping objective, while Stage II generates the full continuous action sequence conditioned on these keyframes. Furthermore, to mitigate cumulative drift, we keep the gripper's orientation being learned while post-optimizing its translation based on the grasping heuristic and kinematic consistency. In both simulation and real-robot experiments, our framework enables stable and precise hand-to-gripper transfer of complex spatial manipulations, significantly outperforming traditional baselines. Project page: https://cosmoh2g.github.io.