WorldDiT: 세계 및 행동 모델링을 위한 통합 확산 아키텍처
WorldDiT: A Unified Diffusion Architecture for World and Action Modeling
July 27, 2026
저자: Sen Wang, R. Gnana Praveen, Bidhan Roy, Marcos Villagra
cs.AI
초록
최근 많은 로봇 정책들은 대규모 사전 학습된 시각-언어 모델(VLM)을 행동 백본(backbone)으로 활용하여 더 강력한 제어를 추구하고 있다. 본 연구에서는 행동 생성과 시각적 세계 모델링을 결합하는 통합 확산 트랜스포머(transformer) 아키텍처인 WorldDiT를 제안하며, 대규모 사전 학습된 VLM 행동 백본 없이도 강력한 성능을 달성한다. 훈련 단계에서 단일 확산 트랜스포머는 연속적인 행동 청크(chunk)를 생성하고, 미래 카메라 프레임으로부터 정규화된 RGB 패치 타겟을 예측한다. 네 가지 LIBERO 시뮬레이션 제품군에 걸쳐 WorldDiT는 네 가지 제품군 모두를 보고하는 방법론들 중 전체 모델 파라미터와 평균 성공률 측면에서 보고된 파레토 프론티어(Pareto frontier)에 위치한다. 이러한 결과는 향후 규모 확장 연구를 위한 강력한 10억 파라미터 미만의 기준선(baseline)을 제공한다.
English
Many recent robot policies pursue stronger control by using large pretrained vision-language models (VLMs) as the action backbone. We introduce WorldDiT, a unified diffusion transformer architecture that couples action generation with visual world modeling and achieves strong performance without a large pretrained VLM action backbone. During training, a single diffusion transformer generates continuous action chunks and predicts normalized RGB patch targets from future camera frames. Across four LIBERO simulation suites, WorldDiT lies on the reported Pareto frontier for total model parameters and mean success among methods reporting all four suites. These results provide a strong sub-billion-parameter baseline for future scaling studies.