Matrix-Game 3.5: パッチメモリによるリアルタイムストリーミング・インタラクティブワールドモデルの強化
Matrix-Game 3.5: Enhancing Real-Time Streaming Interactive World Models with Patch Memory
August 30, 2026
著者: Runjia Qian, Zile Wang, Jihai Zhang, Kai Zou, Wei Yu, Jiaxing Li, Zexiang Liu, Yaokun Li, Fei Kang, Kaichen Huang, Mengyin An, Haobo Zhang, Biao Jiang, Jiahua Wang, Haofeng Sun, Yang Liu, Yangguang Li
cs.AI
要旨
インタラクティブな世界モデルは、ビデオ生成を、オフラインクリップ合成からインタラクティブな仮想世界の持続的シミュレーションへと拡張し、ゲーム、ロボティクス、身体化エージェント、XRなどの応用を可能にする。しかし、安定した長期的なインタラクティブ生成の実現は依然として困難であり、モデルはリアルタイム自己回帰生成を支えつつ、シーンの幾何構造、動的整合性、カメラ制御を同時に維持しなければならない。本稿では、Matrix-Game 3.0を基盤として、図1に示すMatrix-Game 3.5を提案する。本手法は、以下の3つの主要な改良により、リアルタイムのインタラクティブワールド生成を、幾何構造を考慮し、長期的に整合性のあるシミュレーションへと前進させる。第1に、統合的幾何構造認識メモリフレームワークを提案する。そのパッチメモリとTiled-PRoPEは追加の学習可能パラメータを一切導入せず、明示的な3Dパッチ検索と射影カメラ条件付けを組み合わせることで、幾何学的に整合的なカメラ制御と忠実な長期的シーン想起を実現する。第2に、静的なシーンの幾何構造と動的被写体を別々にモデル化する静的・動的分離型世界表現を導入し、長期的生成の全体を通して幾何的一貫性と被写体の同一性の両方を維持する。第3に、双方向拡散モデルを、Perceptual Flow MatchingとカリキュラムベースのSelf-Rollout DMDによって数ステップの因果的生成器へ変換する、2段階の漸進的リアルタイム蒸留フレームワークを開発し、分単位でのリアルタイムインタラクティブ生成を可能にする。広範な実験により、Unrealシミュレーション環境、オープンワールドゲーム、インターネット動画にわたる統一的学習コーパスを用いることで、MatrixGame 3.5が、長期的シーン想起、精密なカメラ制御、被写体の一貫性、プロンプト駆動のワールド生成、および安定したリアルタイム・オープンワールドインタラクションにおいて優れた性能を達成することが実証された。
English
Interactive world models extend video generation from offline clip synthesis toward persistent simulation of interactive virtual worlds, enabling applications in games, robotics, embodied agents, and XR. Achieving stable long-horizon interactive generation, however, remains challenging, as the model must simultaneously preserve scene geometry, dynamic consistency, and camera control while supporting real-time autoregressive generation. Building upon Matrix-Game 3.0, we present Matrix-Game 3.5, as shown in Figure 1, which advances real-time interactive world generation toward geometry-aware and long-horizon consistent simulation through three key improvements. First, we propose a unified geometry-aware memory framework, whose patch-memory and tiled-PRoPE components introduce no additional learnable parameters, combining explicit 3D patch retrieval with projective camera conditioning to enable geometry-consistent camera control and faithful long-horizon scene recall. Second, we introduce a static-dynamic disentangled world representation that separately models static scene geometry and dynamic subjects, preserving both geometric consistency and subject identity throughout long-horizon generation. Third, we develop a two-stage progressive real-time distillation framework that converts a bidirectional diffusion model into a few-step causal generator through Perceptual Flow Matching and curriculum based Self-Rollout DMD, enabling minute-long real-time interactive generation. Extensive experiments demonstrate that, with a unified training corpus spanning Unreal simulation environments, open-world games, and internet videos, MatrixGame 3.5 achieves strong performance in long-horizon scene recall, precise camera control, subject consistency, prompt-driven world generation, and stable real-time open-world interaction.