ChatPaper.aiChatPaper

RynnWorld-Teleop: 디지털 원격 조작을 위한 행동 조건부 세계 모델

RynnWorld-Teleop: An Action-Conditioned World Model for Digital Teleoperation

July 7, 2026
저자: Haoyu Zhao, Xingyue Zhao, Hangyu Li, Biao Gong, Kehan Li, Siteng Huang, Xin Li, Deli Zhao, Zhongyu Li
cs.AI

초록

로봇 학습의 확장에는 방대하고 다양한 궤적 데이터가 필요하지만, 현재 데이터 수집은 물리적 원격 조작(teleoperation)에 의해 병목 현상이 발생하고 있으며, 이 경우 모든 시연(demonstration)이 작업자의 시간을 특정 하드웨어와 작업 공간에 묶어 놓습니다. 우리는 디지털 원격 조작(digital teleoperation)이라는 패러다임을 소개합니다. 이는 실제 로봇을 생성적 세계 모델(generative world model)로 대체함으로써 데이터 수집을 물리적 제약으로부터 분리합니다. 이 프레임워크에서 작업자의 손 자세( hand-pose) 스트림은 로봇 중심의 생성적 세계 모델을 구동하여 단일 참조 이미지로부터 고충실도 1인칭 비디오를 합성합니다. 기록된 자세 스트림은 구현 방식에 구애받지 않는(embodiment-agnostic) 행동 레이블 역할을 하며, 표준 리타겟팅(retargeting)을 통해 모든 대상 로봇으로 전송 가능하여 물리적 하드웨어와 무관한 완전한 상태-행동 궤적을 모방 학습에 제공합니다. 우리는 이 패러다임을 RynnWorld-Teleop 시스템으로 구현합니다. 이 시스템은 깊이 인식 골격 조건화(depth-aware skeletal conditioning), 비디오 확산 트랜스포머(Diffusion Transformer)에 대한 점진적 인간-로봇 훈련, 그리고 스트리밍 자동회귀 증류(streaming autoregressive distillation)를 통합합니다. 이 파이프라인은 생성 과정을 단일 패스 추론으로 압축하여 단일 H100 GPU에서 40FPS 이상의 실시간 대화형 생성을 가능하게 합니다. RynnWorld-Teleop이 생성한 데이터만으로 훈련된 정책(policy)은 정교하고 다양한 양손 작업에서 효과적인 제로샷 Sim2Real 전이(zero-shot Sim2Real transfer)를 달성합니다. 또한 실제 세계 데이터셋을 당사의 디지털 원격 조작 데이터로 증강하면 성공률이 지속적으로 향상되어, RynnWorld-Teleop이 차세대 로봇 에이전트를 위한 고충실도이자 확장 가능한 데이터 엔진 역할을 함을 입증합니다.
English
Scaling robot learning requires massive, diverse trajectory data, yet collection is currently bottlenecked by physical teleoperation, where every demonstration binds operator time to specific hardware and workspaces. We introduce digital teleoperation, a paradigm that decouples data collection from physical constraints by replacing the real robot with a generative world model. In this framework, an operator's hand-pose stream drives a robot-centric generative world model to synthesize high-fidelity egocentric videos from a single reference image. The recorded pose stream serves as an embodiment-agnostic action label transferable to any target robot via standard retargeting, yielding complete state-action trajectories for imitation learning independent of physical hardware. We instantiate this paradigm in RynnWorld-Teleop, a system that integrates depth-aware skeletal conditioning, progressive human-to-robot training on a video Diffusion Transformer, and streaming autoregressive distillation. This pipeline compresses the generative process into a single-pass inference, enabling 40+ FPS, real-time interactive generation on a single H100 GPU. Policies trained exclusively on RynnWorld-Teleop-generated data achieve effective zero-shot Sim2Real transfer across dexterous and diverse bimanual tasks. Moreover, augmenting real-world datasets with our digitally teleoperated data consistently improves success rates, demonstrating that RynnWorld-Teleop serves as a high-fidelity, scalable data engine for the next generation of robotic agents.