WorldDiT:統一擴散架構於世界與動作建模
WorldDiT: A Unified Diffusion Architecture for World and Action Modeling
July 27, 2026
作者: Sen Wang, R. Gnana Praveen, Bidhan Roy, Marcos Villagra
cs.AI
摘要
近期許多機器人政策試圖透過採用大型預訓練視覺語言模型作為動作主幹,以實現更強的控制能力。我們提出世界擴散變換器(WorldDiT),這是一個統一的擴散變換器架構,將動作生成與視覺世界建模結合,在無需大型預訓練視覺語言模型動作主幹的情況下即展現出優異表現。在訓練過程中,單一擴散變換器會生成連續動作區塊,並從未來攝影機影像中預測正規化的RGB區塊目標。在四個LIBERO模擬套件中,WorldDiT在所有四個套件皆有報告的方法中,其總模型參數與平均成功率均位於所報導的帕累托前沿上。這些結果為未來規模擴展研究提供了強而有力的十億參數以下基準。
English
Many recent robot policies pursue stronger control by using large pretrained vision-language models (VLMs) as the action backbone. We introduce WorldDiT, a unified diffusion transformer architecture that couples action generation with visual world modeling and achieves strong performance without a large pretrained VLM action backbone. During training, a single diffusion transformer generates continuous action chunks and predicts normalized RGB patch targets from future camera frames. Across four LIBERO simulation suites, WorldDiT lies on the reported Pareto frontier for total model parameters and mean success among methods reporting all four suites. These results provide a strong sub-billion-parameter baseline for future scaling studies.