ChatPaper.aiChatPaper

RynnWorld-4D: 로봇 조작을 위한 4D 체화 세계 모델

RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation

July 7, 2026
저자: Haoyu Zhao, Xingyue Zhao, Siteng Huang, Xin Li, Deli Zhao, Zhongyu Li
cs.AI

초록

오픈 월드에서의 로봇 조작은 장면이 어떻게 보이는지 인식하는 것뿐만 아니라, 상호작용 하에서 3차원 구조가 어떻게 움직일지 예측하는 것을 요구한다. 본 논문에서는 동기화된 RGB, 깊이, 옵티컬 플로우, 즉 RGB-DF가 장면의 근본적인 4D 동역학을 포착하는 물리적으로 기반한 표현을 제공한다고 주장한다. 2D 픽셀 비디오와 비교하여, 이러한 멀티모달 시너지는 시각적 외형을 기하학적 구조 및 시간적 움직임과 정렬시켜, 로봇 시스템이 요구하는 저수준 엔드 이펙터 동작에 훨씬 가까운 표현 공간을 생성함으로써 세계 예측과 정책 학습 간의 격차를 좁힌다. 이러한 통찰을 바탕으로, 단일 RGB-D 이미지와 언어 명령으로부터 미래의 RGB 프레임, 깊이 맵, 옵티컬 플로우를 하나의 통합된 확산 과정 내에서 공동 생성하는 생성 모델인 RynnWorld-4D를 소개한다. 이 4D 월드 모델은 크로스모달 어텐션을 프레임별 3D RoPE와 통합하는 삼중 분기 아키텍처를 특징으로 하여, 외형, 기하학, 움직임이 일관되게 진화하도록 보장한다. 대규모 훈련 데이터를 공급하기 위해, 우리는 2억 5440만 프레임 이상의 자아중심 인간 및 로봇 조작 비디오에 깊이와 옵티컬 플로우에 대한 고품질 의사 레이블을 갖춘 대규모 데이터셋인 Rynn4DDataset 1.0을 구축했다. 또한, RynnWorld-4D의 내부 4D 표현을 단일 순방향 패스에서 소비하여 비용이 많이 드는 다단계 노이즈 제거를 우회하고, 폐쇄 루프 방식으로 로봇 동작을 출력하는 역동역학 헤드인 RynnWorld-4D-Policy를 제안한다. 실험 결과, RynnWorld-4D는 시간적 및 공간적으로 일관된 4D 예측을 생성하며, RynnWorld-4D-Policy는 실제 세계의 정밀한 양손 조작 작업, 특히 공간적 정밀도와 시간적 조정을 요구하는 작업에서 최첨단 성능을 달성함을 보여준다.
English
Robotic manipulation in the open world requires not only recognizing what a scene looks like, but also anticipating how its 3D structure moves under interaction. We argue that synchronized RGB, depth, and optical flow, namely RGB-DF, provide a physically grounded representation that captures the underlying 4D dynamics of a scene. Compared to 2D pixel videos, this multi-modal synergy aligns visual appearance with geometric structure and temporal motion, creating a representation space significantly closer to the low-level end-effector actions demanded by robotic systems, thereby narrowing the gap between world prediction and policy learning. Building on this insight, we introduce RynnWorld-4D, a generative model that co-produces future RGB frames, depth maps, and optical flow from a single RGB-D image and a language instruction within one unified diffusion process. This 4D world model features a tri-branch architecture that integrates cross-modal attention with frame-wise 3D RoPE, ensuring that appearance, geometry, and motion evolve consistently. To supply training data at scale, we curate Rynn4DDataset 1.0, a massive dataset of over 254.4 million frames across egocentric human and robotic manipulation videos with high-quality pseudo-labels for depth and optical flow. We further propose RynnWorld-4D-Policy, an inverse dynamics head that consumes the internal 4D representations of RynnWorld-4D in a single forward pass, bypassing expensive multi-step denoising, to output robot actions in a closed-loop manner. Experiments show that RynnWorld-4D produces temporally and spatially coherent 4D predictions, and that RynnWorld-4D-Policy achieves state-of-the-art performance on real-world dexterous bimanual manipulation tasks, particularly excelling in tasks demanding spatial precision and temporal coordination.