Matrix-Game 3.5:利用區塊記憶增強即時串流互動式世界模型
Matrix-Game 3.5: Enhancing Real-Time Streaming Interactive World Models with Patch Memory
August 30, 2026
作者: Runjia Qian, Zile Wang, Jihai Zhang, Kai Zou, Wei Yu, Jiaxing Li, Zexiang Liu, Yaokun Li, Fei Kang, Kaichen Huang, Mengyin An, Haobo Zhang, Biao Jiang, Jiahua Wang, Haofeng Sun, Yang Liu, Yangguang Li
cs.AI
摘要
互動式世界模型將影片生成從離線片段合成擴展至對互動式虛擬世界的持久模擬,使其能應用於遊戲、機器人、具身智能體與擴展現實等領域。然而,實現穩定的長時域互動生成仍具挑戰性,因為模型必須同時保持場景幾何、動態一致性與相機控制,並支援即時自迴歸生成。基於Matrix-Game 3.0,我們提出Matrix-Game 3.5,如圖1所示,透過三項關鍵改進,將即時互動式世界生成推進至幾何感知與長時域一致的模擬。首先,我們提出統一的幾何感知記憶框架,其區塊記憶與分塊PRoPE組件不引入額外可學習參數,結合顯式3D區塊檢索與投影相機條件約束,實現幾何一致的相機控制與忠實的長時域場景回溯。其次,我們引入靜態-動態解耦的世界表徵,分別建模靜態場景幾何與動態主體,從而在長時域生成中同時保持幾何一致性與主體身份。第三,我們開發了兩階段漸進式即時蒸餾框架,透過感知流匹配與基於課程學習的自滾動DMD,將雙向擴散模型轉化為少步驟因果生成器,實現長達數分鐘的即時互動生成。大量實驗表明,憑藉涵蓋Unreal模擬環境、開放世界遊戲與網路影片的統一訓練語料庫,Matrix-Game 3.5在長時域場景回溯、精確相機控制、主體一致性、提示驅動的世界生成以及穩定的即時開放世界互動方面均展現出優異性能。
English
Interactive world models extend video generation from offline clip synthesis toward persistent simulation of interactive virtual worlds, enabling applications in games, robotics, embodied agents, and XR. Achieving stable long-horizon interactive generation, however, remains challenging, as the model must simultaneously preserve scene geometry, dynamic consistency, and camera control while supporting real-time autoregressive generation. Building upon Matrix-Game 3.0, we present Matrix-Game 3.5, as shown in Figure 1, which advances real-time interactive world generation toward geometry-aware and long-horizon consistent simulation through three key improvements. First, we propose a unified geometry-aware memory framework, whose patch-memory and tiled-PRoPE components introduce no additional learnable parameters, combining explicit 3D patch retrieval with projective camera conditioning to enable geometry-consistent camera control and faithful long-horizon scene recall. Second, we introduce a static-dynamic disentangled world representation that separately models static scene geometry and dynamic subjects, preserving both geometric consistency and subject identity throughout long-horizon generation. Third, we develop a two-stage progressive real-time distillation framework that converts a bidirectional diffusion model into a few-step causal generator through Perceptual Flow Matching and curriculum based Self-Rollout DMD, enabling minute-long real-time interactive generation. Extensive experiments demonstrate that, with a unified training corpus spanning Unreal simulation environments, open-world games, and internet videos, MatrixGame 3.5 achieves strong performance in long-horizon scene recall, precise camera control, subject consistency, prompt-driven world generation, and stable real-time open-world interaction.