RynnWorld-Teleop: デジタル遠隔操作のための行動条件付き世界モデル
RynnWorld-Teleop: An Action-Conditioned World Model for Digital Teleoperation
July 7, 2026
著者: Haoyu Zhao, Xingyue Zhao, Hangyu Li, Biao Gong, Kehan Li, Siteng Huang, Xin Li, Deli Zhao, Zhongyu Li
cs.AI
要旨
ロボット学習のスケールアップには、多様で大規模な軌跡データが必要不可欠である。しかし現状、データ収集は物理的なテレオペレーションに制約されており、すべての実演がオペレーターの時間を特定のハードウェアと作業空間に縛り付けている。本稿では、実ロボットを生成的世界モデルで置き換えることにより、データ収集を物理的制約から切り離す新たなパラダイム「デジタルテレオペレーション」を紹介する。このフレームワークでは、オペレーターの手姿勢ストリームがロボット中心の生成的世界モデルを駆動し、単一の参照画像から高忠実度の一人称視点ビデオを合成する。記録された姿勢ストリームは、任意の対象ロボットに標準的なリターゲティングを介して転送可能な、身体に依存しない行動ラベルとして機能し、物理的なハードウェアに依存しない模倣学習のための完全な状態-行動軌跡を生成する。我々はこのパラダイムを、深度認識スケルトン条件付け、プログレッシブな人間からロボットへの訓練(ビデオ拡散トランスフォーマー上)、およびストリーミング自己回帰蒸留を統合したシステム「RynnWorld-Teleop」として具体化した。このパイプラインは生成プロセスを単一パスの推論に圧縮し、単一のH100 GPU上で毎秒40フレーム以上のリアルタイムインタラクティブ生成を可能にする。RynnWorld-Teleopが生成したデータのみで訓練されたポリシーは、器用かつ多様な両手操作タスクにおいて、効果的なゼロショットSim2Real転送を達成する。さらに、実世界データセットを本デジタルテレオペレーションデータで拡張することにより、成功率が一貫して向上することを示し、RynnWorld-Teleopが次世代ロボットエージェントのための高忠実度かつスケーラブルなデータエンジンとして機能することを実証する。
English
Scaling robot learning requires massive, diverse trajectory data, yet collection is currently bottlenecked by physical teleoperation, where every demonstration binds operator time to specific hardware and workspaces. We introduce digital teleoperation, a paradigm that decouples data collection from physical constraints by replacing the real robot with a generative world model. In this framework, an operator's hand-pose stream drives a robot-centric generative world model to synthesize high-fidelity egocentric videos from a single reference image. The recorded pose stream serves as an embodiment-agnostic action label transferable to any target robot via standard retargeting, yielding complete state-action trajectories for imitation learning independent of physical hardware. We instantiate this paradigm in RynnWorld-Teleop, a system that integrates depth-aware skeletal conditioning, progressive human-to-robot training on a video Diffusion Transformer, and streaming autoregressive distillation. This pipeline compresses the generative process into a single-pass inference, enabling 40+ FPS, real-time interactive generation on a single H100 GPU. Policies trained exclusively on RynnWorld-Teleop-generated data achieve effective zero-shot Sim2Real transfer across dexterous and diverse bimanual tasks. Moreover, augmenting real-world datasets with our digitally teleoperated data consistently improves success rates, demonstrating that RynnWorld-Teleop serves as a high-fidelity, scalable data engine for the next generation of robotic agents.