ChatPaper.aiChatPaper

RynnWorld-4D: 用于机器人操作的4D具身世界模型

RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation

July 7, 2026
作者: Haoyu Zhao, Xingyue Zhao, Siteng Huang, Xin Li, Deli Zhao, Zhongyu Li
cs.AI

摘要

在开放世界中进行机器人操作,不仅需要识别场景的视觉外观,还需预测其三维结构在交互过程中如何变化。我们认为,同步的RGB、深度与光流信息(即RGB-DF)能够提供一种物理上的有根基的表示,捕捉场景背后的四维动态。相较于二维像素视频,这种多模态融合将视觉外观、几何结构与时序运动对齐,构建出一个更接近机器人系统所需低级末端执行器动作的表示空间,从而缩小了世界预测与策略学习之间的差距。基于这一洞察,我们提出了RynnWorld-4D——一个生成式模型,它能够在统一的扩散过程中,从单张RGB-D图像与一条语言指令出发,共同生成未来的RGB帧、深度图与光流。该4D世界模型采用三分支架构,集成跨模态注意力与逐帧三维旋转位置编码(3D RoPE),确保外观、几何与运动的一致性演化。为提供大规模训练数据,我们整理了Rynn4DDataset 1.0——一个包含超过2.544亿帧的大规模数据集,涵盖自我中心视角的人类与机器人操作视频,并附有高质量的深度与光流伪标签。我们进一步提出了RynnWorld-4D-Policy——一个逆动力学头,能够通过单次前向传播直接消费RynnWorld-4D的内部4D表示,绕过耗时的多步去噪过程,以闭环方式输出机器人动作。实验表明,RynnWorld-4D能够生成时空一致的4D预测,而RynnWorld-4D-Policy在真实世界的灵巧双臂操作任务中达到了最先进的性能,尤其在需要空间精度与时间协调的任务中表现卓越。
English
Robotic manipulation in the open world requires not only recognizing what a scene looks like, but also anticipating how its 3D structure moves under interaction. We argue that synchronized RGB, depth, and optical flow, namely RGB-DF, provide a physically grounded representation that captures the underlying 4D dynamics of a scene. Compared to 2D pixel videos, this multi-modal synergy aligns visual appearance with geometric structure and temporal motion, creating a representation space significantly closer to the low-level end-effector actions demanded by robotic systems, thereby narrowing the gap between world prediction and policy learning. Building on this insight, we introduce RynnWorld-4D, a generative model that co-produces future RGB frames, depth maps, and optical flow from a single RGB-D image and a language instruction within one unified diffusion process. This 4D world model features a tri-branch architecture that integrates cross-modal attention with frame-wise 3D RoPE, ensuring that appearance, geometry, and motion evolve consistently. To supply training data at scale, we curate Rynn4DDataset 1.0, a massive dataset of over 254.4 million frames across egocentric human and robotic manipulation videos with high-quality pseudo-labels for depth and optical flow. We further propose RynnWorld-4D-Policy, an inverse dynamics head that consumes the internal 4D representations of RynnWorld-4D in a single forward pass, bypassing expensive multi-step denoising, to output robot actions in a closed-loop manner. Experiments show that RynnWorld-4D produces temporally and spatially coherent 4D predictions, and that RynnWorld-4D-Policy achieves state-of-the-art performance on real-world dexterous bimanual manipulation tasks, particularly excelling in tasks demanding spatial precision and temporal coordination.