ChatPaper.aiChatPaper

帶有世界狀態暫存器的串流多智能體自迴歸擴散模型

Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers

July 23, 2026
作者: Sicheng Mo, Yuheng Li, Ziyang Leng, Krishna Kumar Singh, Bolei Zhou
cs.AI

摘要

多智能體互動式世界模型不僅應生成一致的觀測結果,還需維持跨智能體持久存在並隨視角演化的世界狀態。現有的自回歸影片擴散管線將觀測歷史作為條件上下文傳遞,這使得共享狀態在多重智能體與多重視角設定下難以維持。我們提出WorldWeaver(W^2),一種串流多智能體影片擴散模型,該模型在推論過程中引入跨智能體世界狀態暫存器:可學習的標記(token)用以儲存共享世界資訊、追蹤個別智能體狀態,並在每個生成區塊後動態更新。我們透過涵蓋個別智能體狀態、全局視角狀態(包括鳥瞰圖)及場景文本的監督訊號來奠基這些暫存器。此外,我們以混合Transformer(Mixture-of-Transformers)設計改善架構,分別使用獨立權重進行世界狀態建模與視覺畫面建模。在雙智能體Minecraft影片生成的大量實驗中顯示,明確的世界狀態建模能提升邏輯一致性與生成品質。
English
Multi-agent interactive world models should not only generate consistent observations, but also maintain world states that persist across agents and evolve across views. Existing autoregressive video diffusion pipelines carry forward observation history as conditioning context, which makes shared state difficult to maintain in multi-agent and multi-view settings. We present WorldWeaver (W^2), a streaming multi-agent video diffusion model that augments rollout with cross-agent world state registers: learnable tokens that store shared world information, track individual agent status, and are dynamically updated after each generated chunk. We ground these registers with supervision signals spanning individual agent status, global state views including bird's-eye views, and scene text. We further improve the architecture with a Mixture-of-Transformers design that uses separate weights for world state modeling and visual frame modeling. Extensive experiments in two-agent Minecraft video generation show that explicit world-state modeling improves logical consistency and generation quality.