表現オートエンコーダを用いたマルチプレイヤーインタラクティブ世界モデル
Multiplayer Interactive World Models with Representation Autoencoders
July 6, 2026
著者: Anthony Hu, Václav Volhejn, Adrien Ramanana Rahary, Chris Mulder, Aditya Makkar, Amélie Royer, Manu Orsini, Alyx Liao, Adam Jelley, Eloi Alonso, Florian Laurent, Fredrik Norén, James Swingos, Jan Hünermann, Kent Rollins, Lucas Hosseini, Matthieu Le Cauchois, Maxim Peter, Pim de Witte, Tim Brown, Vincent Micheli, Moritz Böhle, Gabriel de Marmiesse, Viktoriia Sharmanska, Lucia Specia, Michael Black, Patrick Pérez
cs.AI
要旨
本論文では、複雑な物理的相互作用によって支配される高度に動的な環境を対象とした、初のマルチプレイヤー世界モデルを提案する。シングルプレイヤー世界モデルが他のエージェントを環境の一部として扱うのに対し、本モデルは複数エージェントの行動系列に条件付けを行い、場面の変化を正しいプレイヤーに帰属させる学習と、任意の行動の組み合わせ下での一貫性の維持を実現する。我々は、高速で密結合なダイナミクスのもとでプレイヤーが競争と協力を行うゲーム「ロケットリーグ」においてこの問題を研究した。公開ボットを用いて収集した1万時間のゲームプレイデータで学習した50億パラメータの潜在拡散モデルは、単一のNVIDIA B200 GPU上で毎秒20フレームを生成し、4人プレイのマッチをリアルタイムで再現する。短いクリップでのみ学習されているにもかかわらず、そのロールアウトは学習時系列をはるかに超えて安定する。分布品質は測定した最長の5分間にわたって安定して維持され、実際には数時間にわたって崩壊の兆候なくロールアウトが継続することを確認している。我々は、ビデオコーデック、生成目的関数、マルチプレイヤー条件付け方式という中心的設計選択について体系的に調査する。さらに、モデルとデータのスケールに応じて行動がどのように変化するかを、創発する能力と残存する障害モードを含めて特徴づける。また、視覚的外観のみならず物理的理解を評価する、標的を絞った評価手法を開発する。マルチプレイヤー世界モデルに関する研究の継続を支援するため、データセット、学習・推論コードベースの全容、およびライブデモを公開する。
English
We introduce the first multiplayer world model for highly dynamic environments governed by complex physical interactions. Whereas single-player world models treat the other agents as part of the environment, ours conditions on the action streams of multiple agents, learning to attribute changes in the scene to the correct player and to stay coherent under arbitrary combinations of their actions. We study this problem in the game of Rocket League, where players compete and cooperate under fast, tightly coupled dynamics. Trained on 10,000 hours of gameplay collected with publicly available bots, our 5-billion-parameter latent diffusion model generates four-player matches in real time, producing 20 frames per second on a single Nvidia B200 GPU. Although trained only on short clips, its rollouts stay stable far beyond the training horizon: distributional quality holds steady out to five minutes, the longest horizon we measure, and in practice we observe rollouts continuing for hours with no sign of collapse. We systematically investigate the central design choices: the video codec, the generative objective, and the multiplayer conditioning scheme. In addition, we characterize how behavior changes with model and data scale, including the capabilities that emerge and the failure modes that persist. We further develop targeted evaluations that probe the model's physical understanding rather than visual appearance alone. To support continued research on multiplayer world models, we release our dataset, our full training and inference codebase, and a live demo.