CosmoH2G: 복잡한 공간 움직임을 수반하는 객체 조작을 위한 손-그리퍼 전이 데이터셋 및 베이스라인 방법
CosmoH2G: A Hand-to-Gripper Transfer Dataset and Baseline Method for Object Manipulation with Complex Spatial Movements
September 7, 2026
저자: Hongxiang Zhao, Mutian Xu, Zeyu Jin, Yiming Hao, Shuguang Cui, Xiaoguang Han
cs.AI
초록
인간 손 시연을 로봇 그리퍼로 전이하는 것은 최근 로봇 학습을 위한 비용 효율적인 해법으로 부상했다. 그러나 기존 방법들은 대부분 단순한 평면 작업에 국한되어 있으며, 로봇 조작에 필수적인 복잡한 공간적 움직임(예: 회전이나 뒤집기를 포함하는 복잡한 궤적)을 처리하지 못한다. 이러한 격차에 착안하여, 우리는 세밀한 손 자세 움직임에 의해 안내되는 암시적·데이터 기반 접근법을 채택한다. 이를 위해 우리는 움직임 복잡도를 우선시하고 원활한 동작 모방을 위해 핸드헬드 그리퍼를 활용하는 엄격한 프로토콜에 따라 손-그리퍼 쌍 시연을 수집하는 확장 가능한 수집 파이프라인을 도입한다. 이는 1,254개의 고유 객체에 걸친 6,189개의 에피소드로 구성된 대규모 쌍 데이터셋을 산출하며, 기존 벤치마크보다 현저히 높은 공간적 복잡성을 보인다. 그러나 이러한 복잡한 매핑을 학습하는 것은 여전히 어렵다. 우리는 전체 그리퍼 포즈 시퀀스를 단순한 엔드투엔드 방식으로 생성하는 것이 불충분하다는 것을 관찰하는데, 이는 복잡한 동역학 아래에서 작은 궤적 편차가 빠르게 누적되기 때문이다. 이를 해결하기 위해, 우리는 2단계 프레임워크를 제안한다. 1단계에서는 매핑 목표를 단순화하기 위해 희소 그리퍼 키프레임(초기 및 최종)을 예측하고, 2단계에서는 이러한 키프레임을 조건으로 전체 연속 동작 시퀀스를 생성한다. 나아가 누적 드리프트를 완화하기 위해 그리퍼의 방향은 학습되도록 유지하는 한편, 파지 휴리스틱과 운동학적 일관성에 기반하여 그리퍼의 이동을 사후 최적화한다. 시뮬레이션과 실제 로봇 실험 모두에서, 우리의 프레임워크는 복잡한 공간 조작의 안정적이고 정밀한 손-그리퍼 전이를 가능하게 하며, 전통적 베이스라인을 크게 능가한다. 프로젝트 페이지: https://cosmoh2g.github.io.
English
Transferring human hand demonstrations to robotic grippers has recently emerged as a cost-effective solution for robot learning. However, existing methods are largely confined to simple, planar tasks and fail to handle complex spatial movements (e.g., intricate trajectories involving rotations or flips) that are essential for robot manipulation. Motivated by this gap, we adopt an implicit, data-driven approach guided by fine-grained hand-pose motions. To this end, we introduce a scalable acquisition pipeline to collect hand-gripper paired demonstrations, governed by a rigorous protocol that prioritizes motion complexity and leverages a handheld gripper for seamless action mimicry. This yields a large-scale paired dataset comprising 6,189 episodes across 1,254 unique objects, exhibiting significantly higher spatial complexity than existing benchmarks. However, learning such complex mappings remains challenging. We observe that naive end-to-end generation of full gripper pose sequences is insufficient, as minor trajectory deviations compound rapidly under intricate dynamics. To address this, we propose a two-stage framework: Stage I predicts sparse gripper keyframes (initial and terminal) to simplify the mapping objective, while Stage II generates the full continuous action sequence conditioned on these keyframes. Furthermore, to mitigate cumulative drift, we keep the gripper's orientation being learned while post-optimizing its translation based on the grasping heuristic and kinematic consistency. In both simulation and real-robot experiments, our framework enables stable and precise hand-to-gripper transfer of complex spatial manipulations, significantly outperforming traditional baselines. Project page: https://cosmoh2g.github.io.