ChatPaper.aiChatPaper

Overdraagbare Dynamiek-Priors Leren van Actie naar Wereldmodellering

Learning Transferable Dynamics Priors from Action to World Modeling

June 28, 2026
Auteurs: Ze Huang, Jiahui Zhang, Hairuo Liu, Chenxi Zhang, Ran Cheng, Li Zhang
cs.AI

Samenvatting

We onderzoeken actie-geconditioneerde wereldmodellering als een schaalbare manier om overdraagbare dynamische priori's voor robotleren te leren. Door een model vooraf te trainen om te voorspellen hoe acties de visuele scène-evolutie aansturen, legt het resulterende wereldmodel herbruikbare interactiedynamica vast die verder gaat dan uiterlijk-niveau videogeneratie. Concreet trainen we een multi-view interactief basis diffusie wereldmodel, A2World, vooraf op grootschalige robotmanipulatiedata met echte actieannotaties. We valideren de geleerde dynamische priori's vanuit twee complementaire perspectieven. Ten eerste passen we A2World aan tot een taak- of scènespecifieke realistische simulator, A2World-sim, waarvan de lange-termijn rollouts simulator-gebaseerde beleidsevaluatie en schaalbare wat-als-analyse ondersteunen door echte robotrollouts te vervangen door wereldmodelrollouts. Ten tweede passen we A2World, uitgaande van dezelfde voorgetrainde gewichten, aan tot een video-actie gezamenlijk voorspellingsmodel, A2World-policy, dat acties voorspelt onder visuele en instructieconditionering. Experimenten met simulatiebenchmarks en echte robotomgevingen tonen aan dat actie-geconditioneerde wereldmodel-pretraining overdraagbare dynamische priori's oplevert die zowel simulator-centrisch als beleid-centrisch robotleren ten goede komen.
English
We study action-conditioned world modeling as a scalable way to learn transferable dynamics priors for robot learning. By pretraining a model to predict how actions drive visual scene evolution, the resulting world model captures reusable interaction dynamics beyond appearance-level video generation. Concretely, we pretrain a multi-view interactive base diffusion world model, A2World, on large-scale robot manipulation data with real action annotations. We validate the learned dynamics priors from two complementary perspectives. First, we adapt A2World into a task- or scene-specialized real-world simulator, A2World-sim, whose long-horizon rollouts support simulator-based policy evaluation and scalable what-if analysis by replacing real-robot rollouts with world model rollouts. Second, starting from the same pretrained weights, we adapt A2World into a video-action joint prediction model, A2World-policy, that predicts actions under visual and instruction conditioning. Experiments across simulation benchmarks and real-robot settings demonstrate that action-conditioned world model pretraining yields transferable dynamics priors that benefit both simulator-centric and policy-centric robot learning.