CosmoH2G:具複雜空間運動之物體操作的手部至夾爪轉移資料集與基線方法
CosmoH2G: A Hand-to-Gripper Transfer Dataset and Baseline Method for Object Manipulation with Complex Spatial Movements
September 7, 2026
作者: Hongxiang Zhao, Mutian Xu, Zeyu Jin, Yiming Hao, Shuguang Cui, Xiaoguang Han
cs.AI
摘要
將人類手部示範遷移至機器人夾爪,近來已成為機器人學習中具成本效益的解決方案。然而,現有方法大多侷限於簡單的平面任務,無法處理機器人操作不可或缺的複雜空間運動(例如涉及旋轉或翻轉的精細軌跡)。有鑑於此一缺口,我們採用由細粒度手部姿態動作引導的隱式、資料驅動方法。為此,我們提出一套可擴展的資料獲取流程,用以收集手部—夾爪配對示範;此流程遵循嚴謹協定,優先考量動作複雜度,並利用手持式夾爪實現無縫動作模仿。這產生了一個大規模配對資料集,包含橫跨 1,254 個獨特物件的 6,189 個回合,展現出顯著高於現有基準的空間複雜度。然而,學習如此複雜的映射仍具挑戰性。我們觀察到,樸素地端到端生成完整夾爪姿態序列並不足夠,因為微小軌跡偏差會在複雜動力學下迅速累積。為解決此問題,我們提出兩階段框架:第一階段預測稀疏夾爪關鍵幀(初始與終端),以簡化映射目標;第二階段則以這些關鍵幀為條件,生成完整連續動作序列。此外,為減緩累積漂移,我們讓夾爪的朝向持續被學習,同時根據抓取啟發式與運動學一致性對其平移進行後優化。在模擬與真實機器人實驗中,我們的框架能針對複雜空間操作實現穩定且精確的手到夾爪轉移,顯著優於傳統基準方法。專案頁面:https://cosmoh2g.github.io。
English
Transferring human hand demonstrations to robotic grippers has recently emerged as a cost-effective solution for robot learning. However, existing methods are largely confined to simple, planar tasks and fail to handle complex spatial movements (e.g., intricate trajectories involving rotations or flips) that are essential for robot manipulation. Motivated by this gap, we adopt an implicit, data-driven approach guided by fine-grained hand-pose motions. To this end, we introduce a scalable acquisition pipeline to collect hand-gripper paired demonstrations, governed by a rigorous protocol that prioritizes motion complexity and leverages a handheld gripper for seamless action mimicry. This yields a large-scale paired dataset comprising 6,189 episodes across 1,254 unique objects, exhibiting significantly higher spatial complexity than existing benchmarks. However, learning such complex mappings remains challenging. We observe that naive end-to-end generation of full gripper pose sequences is insufficient, as minor trajectory deviations compound rapidly under intricate dynamics. To address this, we propose a two-stage framework: Stage I predicts sparse gripper keyframes (initial and terminal) to simplify the mapping objective, while Stage II generates the full continuous action sequence conditioned on these keyframes. Furthermore, to mitigate cumulative drift, we keep the gripper's orientation being learned while post-optimizing its translation based on the grasping heuristic and kinematic consistency. In both simulation and real-robot experiments, our framework enables stable and precise hand-to-gripper transfer of complex spatial manipulations, significantly outperforming traditional baselines. Project page: https://cosmoh2g.github.io.