ChatPaper.aiChatPaper

RynnWorld-4D: ロボット操作のための4D身体化世界モデル

RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation

July 7, 2026
著者: Haoyu Zhao, Xingyue Zhao, Siteng Huang, Xin Li, Deli Zhao, Zhongyu Li
cs.AI

要旨

オープンワールドにおけるロボット操作は、シーンの見た目を認識するだけでなく、その3次元構造が相互作用の下でどのように動くかを予測する必要がある。我々は、RGB、深度、オプティカルフロー、すなわちRGB-DFが、シーンの根底にある4次元動的構造を捉える物理的に根拠づけられた表現を提供すると主張する。2次元ピクセル動画と比較して、このマルチモーダルな相乗効果は、視覚的外観と幾何学的構造、時間的動作を整合させ、ロボットシステムが要求する低レベルのエンドエフェクタ動作にはるかに近い表現空間を創り出す。これにより、世界予測と方策学習の間のギャップを狭める。この知見に基づき、我々はRynnWorld-4Dを紹介する。これは、単一のRGB-D画像と言語指示から、将来のRGBフレーム、深度マップ、オプティカルフローを、統一された単一の拡散プロセス内で同時生成する生成モデルである。この4次元世界モデルは、三分岐アーキテクチャを備え、クロスモーダルアテンションとフレーム単位の3D RoPEを統合し、外観、幾何、動作が一貫して進化することを保証する。大規模な訓練データを提供するため、我々はRynn4DDataset 1.0を厳選した。これは、自己中心視点の人間及びロボット操作動画にわたり、深度とオプティカルフローの高品質な擬似ラベルを付与した、2億5440万フレーム以上の大規模データセットである。さらに、RynnWorld-4D-Policyを提案する。これは、RynnWorld-4Dの内部4次元表現を単一の順伝播で消費し、高コストな多段階ノイズ除去を回避して、閉ループ方式でロボットのアクションを出力する逆動力学ヘッドである。実験により、RynnWorld-4Dが時間的・空間的に一貫した4次元予測を生成すること、そしてRynnWorld-4D-Policyが現実世界の巧緻な両腕操作タスクにおいて最先端の性能を達成し、特に空間的精度と時間的協調を要するタスクで優れていることを示す。
English
Robotic manipulation in the open world requires not only recognizing what a scene looks like, but also anticipating how its 3D structure moves under interaction. We argue that synchronized RGB, depth, and optical flow, namely RGB-DF, provide a physically grounded representation that captures the underlying 4D dynamics of a scene. Compared to 2D pixel videos, this multi-modal synergy aligns visual appearance with geometric structure and temporal motion, creating a representation space significantly closer to the low-level end-effector actions demanded by robotic systems, thereby narrowing the gap between world prediction and policy learning. Building on this insight, we introduce RynnWorld-4D, a generative model that co-produces future RGB frames, depth maps, and optical flow from a single RGB-D image and a language instruction within one unified diffusion process. This 4D world model features a tri-branch architecture that integrates cross-modal attention with frame-wise 3D RoPE, ensuring that appearance, geometry, and motion evolve consistently. To supply training data at scale, we curate Rynn4DDataset 1.0, a massive dataset of over 254.4 million frames across egocentric human and robotic manipulation videos with high-quality pseudo-labels for depth and optical flow. We further propose RynnWorld-4D-Policy, an inverse dynamics head that consumes the internal 4D representations of RynnWorld-4D in a single forward pass, bypassing expensive multi-step denoising, to output robot actions in a closed-loop manner. Experiments show that RynnWorld-4D produces temporally and spatially coherent 4D predictions, and that RynnWorld-4D-Policy achieves state-of-the-art performance on real-world dexterous bimanual manipulation tasks, particularly excelling in tasks demanding spatial precision and temporal coordination.