Matrix-Game 3.5:利用补丁记忆增强实时流式交互世界模型
Matrix-Game 3.5: Enhancing Real-Time Streaming Interactive World Models with Patch Memory
August 30, 2026
作者: Runjia Qian, Zile Wang, Jihai Zhang, Kai Zou, Wei Yu, Jiaxing Li, Zexiang Liu, Yaokun Li, Fei Kang, Kaichen Huang, Mengyin An, Haobo Zhang, Biao Jiang, Jiahua Wang, Haofeng Sun, Yang Liu, Yangguang Li
cs.AI
摘要
交互式世界模型将视频生成从离线片段合成扩展到交互式虚拟世界的持续仿真,从而支持游戏、机器人、具身智能体和扩展现实(XR)等应用。然而,实现稳定的长时程交互生成仍具挑战性,因为模型必须同时保持场景几何、动态一致性以及相机控制,并支持实时自回归生成。在Matrix-Game 3.0的基础上,我们提出了Matrix-Game 3.5(如图1所示),通过三项关键改进,将实时交互式世界生成推进到几何感知且长时程一致的仿真。首先,我们提出一种统一的几何感知记忆框架,其补丁记忆(patch-memory)和分块式RoPE(tiled-PRoPE)组件不引入额外的可学习参数,将显式3D补丁检索与投影相机条件相结合,从而实现几何一致的相机控制与忠实的长时程场景回溯。其次,我们引入一种静态-动态解耦的世界表示,分别对静态场景几何和动态主体进行建模,从而在长时程生成过程中同时保持几何一致性和主体身份。第三,我们开发了一种两阶段渐进式实时蒸馏框架,通过感知流匹配(Perceptual Flow Matching)和基于课程的自推演分布匹配蒸馏(Self-Rollout DMD),将双向扩散模型转化为少步因果生成器,从而支持分钟级实时交互生成。大量实验表明,凭借涵盖Unreal仿真环境、开放世界游戏和互联网视频的统一训练语料库,MatrixGame 3.5在长时程场景回溯、精确相机控制、主体一致性、提示驱动的世界生成以及稳定实时开放世界交互等方面展现出强大的性能。
English
Interactive world models extend video generation from offline clip synthesis toward persistent simulation of interactive virtual worlds, enabling applications in games, robotics, embodied agents, and XR. Achieving stable long-horizon interactive generation, however, remains challenging, as the model must simultaneously preserve scene geometry, dynamic consistency, and camera control while supporting real-time autoregressive generation. Building upon Matrix-Game 3.0, we present Matrix-Game 3.5, as shown in Figure 1, which advances real-time interactive world generation toward geometry-aware and long-horizon consistent simulation through three key improvements. First, we propose a unified geometry-aware memory framework, whose patch-memory and tiled-PRoPE components introduce no additional learnable parameters, combining explicit 3D patch retrieval with projective camera conditioning to enable geometry-consistent camera control and faithful long-horizon scene recall. Second, we introduce a static-dynamic disentangled world representation that separately models static scene geometry and dynamic subjects, preserving both geometric consistency and subject identity throughout long-horizon generation. Third, we develop a two-stage progressive real-time distillation framework that converts a bidirectional diffusion model into a few-step causal generator through Perceptual Flow Matching and curriculum based Self-Rollout DMD, enabling minute-long real-time interactive generation. Extensive experiments demonstrate that, with a unified training corpus spanning Unreal simulation environments, open-world games, and internet videos, MatrixGame 3.5 achieves strong performance in long-horizon scene recall, precise camera control, subject consistency, prompt-driven world generation, and stable real-time open-world interaction.