ChatPaper.aiChatPaper

使用表徵自編碼器的多人互動世界模型

Multiplayer Interactive World Models with Representation Autoencoders

July 6, 2026
作者: Anthony Hu, Václav Volhejn, Adrien Ramanana Rahary, Chris Mulder, Aditya Makkar, Amélie Royer, Manu Orsini, Alyx Liao, Adam Jelley, Eloi Alonso, Florian Laurent, Fredrik Norén, James Swingos, Jan Hünermann, Kent Rollins, Lucas Hosseini, Matthieu Le Cauchois, Maxim Peter, Pim de Witte, Tim Brown, Vincent Micheli, Moritz Böhle, Gabriel de Marmiesse, Viktoriia Sharmanska, Lucia Specia, Michael Black, Patrick Pérez
cs.AI

摘要

我們介紹了第一個針對高度動態環境(由複雜物理交互主導)的多玩家世界模型。相較於單玩家世界模型將其他智能體視為環境的一部分,我們的模型以多個智能體的動作流為條件,學習將場景變化歸因於正確的玩家,並在任意動作組合下保持一致性。我們在《火箭聯盟》這款遊戲中研究此問題,玩家在快速且緊密耦合的動態環境中競爭與合作。利用由公開可用的機器人收集的一萬小時遊戲數據進行訓練,我們擁有五十億參數的潛在擴散模型能即時生成四玩家比賽,在單個Nvidia B200 GPU上每秒產生20幀畫面。儘管僅以短片片段訓練,其生成序列在遠超過訓練時長的情況下仍保持穩定:分佈品質在長達五分鐘(我們測量的最長時間)內維持穩定,實際觀察中生成序列可持續數小時且無崩潰跡象。我們系統性地探討了核心設計選擇:影片編解碼器、生成目標以及多玩家條件機制。此外,我們描述了行為如何隨模型與數據規模變化,包括湧現的能力與持續存在的失敗模式。我們進一步開發了針對性的評估,以探測模型的物理理解能力,而非僅限於視覺外觀。為支持多玩家世界模型的持續研究,我們開源了數據集、完整的訓練與推論程式碼庫,以及一個即時演示。
English
We introduce the first multiplayer world model for highly dynamic environments governed by complex physical interactions. Whereas single-player world models treat the other agents as part of the environment, ours conditions on the action streams of multiple agents, learning to attribute changes in the scene to the correct player and to stay coherent under arbitrary combinations of their actions. We study this problem in the game of Rocket League, where players compete and cooperate under fast, tightly coupled dynamics. Trained on 10,000 hours of gameplay collected with publicly available bots, our 5-billion-parameter latent diffusion model generates four-player matches in real time, producing 20 frames per second on a single Nvidia B200 GPU. Although trained only on short clips, its rollouts stay stable far beyond the training horizon: distributional quality holds steady out to five minutes, the longest horizon we measure, and in practice we observe rollouts continuing for hours with no sign of collapse. We systematically investigate the central design choices: the video codec, the generative objective, and the multiplayer conditioning scheme. In addition, we characterize how behavior changes with model and data scale, including the capabilities that emerge and the failure modes that persist. We further develop targeted evaluations that probe the model's physical understanding rather than visual appearance alone. To support continued research on multiplayer world models, we release our dataset, our full training and inference codebase, and a live demo.