ChatPaper.aiChatPaper

WorldDiT:一种用于世界与动作建模的统一扩散架构

WorldDiT: A Unified Diffusion Architecture for World and Action Modeling

July 27, 2026
作者: Sen Wang, R. Gnana Praveen, Bidhan Roy, Marcos Villagra
cs.AI

摘要

近年来的许多机器人政策倾向于使用大型预训练视觉语言模型(VLM)作为动作主干,以实现更强的控制能力。为此,我们提出了WorldDiT——一种统一的扩散变换器架构,该架构将动作生成与视觉世界建模相结合,无需大型预训练VLM动作主干即可实现出色性能。在训练过程中,单个扩散变换器能够生成连续的动作片段,并从未来摄像头帧中预测归一化的RGB图像块目标。在四个LIBERO仿真套件中,WorldDiT在报告所有四个套件结果的方法中,其总模型参数与平均成功率均位于帕累托前沿。这些结果为未来的规模化研究提供了强有力的次十亿参数基线。
English
Many recent robot policies pursue stronger control by using large pretrained vision-language models (VLMs) as the action backbone. We introduce WorldDiT, a unified diffusion transformer architecture that couples action generation with visual world modeling and achieves strong performance without a large pretrained VLM action backbone. During training, a single diffusion transformer generates continuous action chunks and predicts normalized RGB patch targets from future camera frames. Across four LIBERO simulation suites, WorldDiT lies on the reported Pareto frontier for total model parameters and mean success among methods reporting all four suites. These results provide a strong sub-billion-parameter baseline for future scaling studies.