RynnWorld-4D: 4D-belichaamde wereldmodellen voor robotmanipulatie
RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation
July 7, 2026
Auteurs: Haoyu Zhao, Xingyue Zhao, Siteng Huang, Xin Li, Deli Zhao, Zhongyu Li
cs.AI
Samenvatting
Robotische manipulatie in de open wereld vereist niet alleen het herkennen van hoe een scène eruitziet, maar ook het anticiperen van hoe de 3D-structuur ervan beweegt onder interactie. Wij stellen dat gesynchroniseerde RGB, diepte en optische stroom, namelijk RGB-DF, een fysiek gefundeerde representatie bieden die de onderliggende 4D-dynamica van een scène vastlegt. In vergelijking met 2D-pixelvideo's brengt deze multimodale synergie visuele verschijning in lijn met geometrische structuur en temporele beweging, waardoor een representatieruimte ontstaat die aanzienlijk dichter bij de laag-niveau eindeffectoracties ligt die robotsystemen vereisen, en zo de kloof tussen wereldvoorspelling en beleidsleren verkleint. Voortbouwend op dit inzicht introduceren wij RynnWorld-4D, een generatief model dat toekomstige RGB-frames, dieptekaarten en optische stroom co-produceert uit een enkel RGB-D-beeld en een taal-instructie binnen één uniform diffusieproces. Dit 4D-wereldmodel heeft een drietakkige architectuur die crossmodale aandacht integreert met framegewijze 3D-RoPE, zodat verschijning, geometrie en beweging consistent evolueren. Om trainingsdata op schaal te leveren, hebben wij Rynn4DDataset 1.0 samengesteld, een enorme dataset van meer dan 254,4 miljoen frames uit egocentrische menselijke en robotische manipulatievideo's met hoogwaardige pseudo-labels voor diepte en optische stroom. Verder stellen wij RynnWorld-4D-Policy voor, een inverse dynamica kop die de interne 4D-representaties van RynnWorld-4D consumeert in één enkele voorwaartse doorgang, waarbij dure meerstapsdenoising wordt omzeild, om robotacties in een gesloten-lus manier uit te voeren. Experimenten tonen aan dat RynnWorld-4D temporeel en ruimtelijk coherente 4D-voorspellingen produceert, en dat RynnWorld-4D-Policy state-of-the-art prestaties levert op realistische behendige bimanuele manipulatietaken, met name uitblinkend in taken die ruimtelijke precisie en temporele coördinatie vereisen.
English
Robotic manipulation in the open world requires not only recognizing what a scene looks like, but also anticipating how its 3D structure moves under interaction. We argue that synchronized RGB, depth, and optical flow, namely RGB-DF, provide a physically grounded representation that captures the underlying 4D dynamics of a scene. Compared to 2D pixel videos, this multi-modal synergy aligns visual appearance with geometric structure and temporal motion, creating a representation space significantly closer to the low-level end-effector actions demanded by robotic systems, thereby narrowing the gap between world prediction and policy learning. Building on this insight, we introduce RynnWorld-4D, a generative model that co-produces future RGB frames, depth maps, and optical flow from a single RGB-D image and a language instruction within one unified diffusion process. This 4D world model features a tri-branch architecture that integrates cross-modal attention with frame-wise 3D RoPE, ensuring that appearance, geometry, and motion evolve consistently. To supply training data at scale, we curate Rynn4DDataset 1.0, a massive dataset of over 254.4 million frames across egocentric human and robotic manipulation videos with high-quality pseudo-labels for depth and optical flow. We further propose RynnWorld-4D-Policy, an inverse dynamics head that consumes the internal 4D representations of RynnWorld-4D in a single forward pass, bypassing expensive multi-step denoising, to output robot actions in a closed-loop manner. Experiments show that RynnWorld-4D produces temporally and spatially coherent 4D predictions, and that RynnWorld-4D-Policy achieves state-of-the-art performance on real-world dexterous bimanual manipulation tasks, particularly excelling in tasks demanding spatial precision and temporal coordination.