ABot-M0.5: Geünificeerd Mobiliteit-en-Manipulatie Wereldactiemodel
ABot-M0.5: Unified Mobility-and-Manipulation World Action Model
July 1, 2026
Auteurs: Ronghan Chen, Yandan Yang, Zuojin Tang, Dongjie Huo, Tong Lin, Haoning Wu, Haoyun Liu, Yuzhi Chen, Lulu Zheng, Botai Yuan, Tianlun Li, Mingxin Wang, Dekang Qi, Bin Hu, Wei Mei, Yuze Xuan, Haolong Yang, Yanqing Zhu, Mu Xu, Zhiheng Ma, Xinyuan Chang
cs.AI
Samenvatting
Mobiele manipulatie is een cruciale vaardigheid voor algemeen inzetbare robots, maar blijft uitdagend voor huidige belichaamde leermethoden. VLA-beleid is typisch reactief en mist expliciete wereldmodellering, terwijl bestaande Wereldactiemodellen (WAM's) nog slecht zijn afgestemd op de structuur van mobiele manipulatie: ze werken op grove videoblokken, modelleren verstrengelde navigatie-manipulatieacties en trainen inverse dynamica onder supervisie die niet overeenkomt met autoregressieve inferentie. Hierdoor missen ze vaak fijnmazige contactdynamica, ondervinden ze conflicten in de actieverdeling en accumuleren ze fouten over langetermijnrollouts. Wij stellen ABot-M0.5 voor, een nieuw WAM gebaseerd op het inzicht dat mobiele manipulatie afstemming vereist op drie niveaus: temporele granulariteit, actieruimte en trainings-testconsistentie. Voor afstemming van temporele granulariteit introduceren we tussenliggende latente acties die lokale visuele toestandsovergangen vastleggen en dienen als een overbruggende actieruimte tussen videolaternten en belichamingsspecifieke besturingen. Voor afstemming van de actieruimte ontwerpen we een dubbele Mixture-of-Transformers-architectuur die zowel modaliteitsrepresentaties als heterogene actiesubruimtes, zoals basisbeweging en armmanipulatie, ontward. Voor afstemming van inferentieomstandigheden stellen we de dream-forcing trainingsstrategie voor, die inverse dynamica progressief traint op modelvoorspelde video's, wat de trainings-testafstemming en robuustheid tijdens autoregressieve voorspelling verbetert. Experimenten op uitdagende mobiele en fijnmazige manipulatiebenchmarks tonen aan dat ABot-M0.5 state-of-the-art prestaties levert, zowel in langdurige taakuitvoering als in nauwkeurige controle. Deze resultaten benadrukken het cruciale belang van granulariteitsafgestemde, actie-ontwarde en inferentie-consistente wereld-actiemodellering.
English
Mobile manipulation is a key capability for general-purpose robots, yet remains challenging for current embodied learning methods. VLA policies are typically reactive and lack explicit world modeling, while existing World Action Models (WAMs) are still poorly aligned with the structure of mobile manipulation: they operate on coarse video chunks, model entangled navigation-manipulation actions, and train inverse dynamics under supervision that does not match autoregressive inference. As a result, they often miss fine-grained contact dynamics, suffer from action-distribution conflicts, and accumulate errors over long-horizon rollouts. We propose ABot-M0.5, a new WAM built on the insight that mobile manipulation requires alignment at three levels: temporal granularity, action space, and train-test consistency. To align temporal granularity, we introduce intermediate latent actions that capture local visual state transitions and serve as an bridging action space between video latents and embodiment-specific controls. To align action space, we design a dual-level Mixture-of-Transformers architecture that disentangles both modality representations and heterogeneous action subspaces such as base movement and arm manipulation. To align inference conditions, we propose the dream-forcing training strategy that progressively trains inverse dynamics on model-predicted videos, improving train-test alignment and robustness during autoregressive prediction. Experiments on challenging mobile and fine-grained manipulation benchmarks demonstrate that ABot-M0.5 achieves state-of-the-art performance in both long-horizon task success and finegrained control accuracy. These results highlight the critical importance of granularity-aligned, action-disentangled, and inference-consistent world-action modeling.