ChatPaper.aiChatPaper

표현 오토인코더를 활용한 멀티플레이어 상호작용 세계 모델

Multiplayer Interactive World Models with Representation Autoencoders

July 6, 2026
저자: Anthony Hu, Václav Volhejn, Adrien Ramanana Rahary, Chris Mulder, Aditya Makkar, Amélie Royer, Manu Orsini, Alyx Liao, Adam Jelley, Eloi Alonso, Florian Laurent, Fredrik Norén, James Swingos, Jan Hünermann, Kent Rollins, Lucas Hosseini, Matthieu Le Cauchois, Maxim Peter, Pim de Witte, Tim Brown, Vincent Micheli, Moritz Böhle, Gabriel de Marmiesse, Viktoriia Sharmanska, Lucia Specia, Michael Black, Patrick Pérez
cs.AI

초록

우리는 복잡한 물리적 상호작용이 지배하는 고도로 동적인 환경을 위한 최초의 멀티플레이어 월드 모델을 소개한다. 싱글플레이어 월드 모델이 다른 에이전트를 환경의 일부로 취급하는 반면, 우리의 모델은 여러 에이전트의 행동 스트림에 조건화되어 장면의 변화를 올바른 플레이어에게 귀속시키고 이들의 행동이 임의로 결합된 상황에서도 일관성을 유지하는 방법을 학습한다. 우리는 플레이어들이 빠르고 긴밀하게 결합된 역학 하에서 경쟁하고 협력하는 게임인 Rocket League에서 이 문제를 연구한다. 공개적으로 이용 가능한 봇을 통해 수집된 10,000시간의 게임플레이로 훈련된 우리의 50억 파라미터 잠재 확산 모델은 실시간으로 4인 매치를 생성하며, 단일 Nvidia B200 GPU에서 초당 20프레임을 생성한다. 짧은 클립으로만 훈련되었지만, 그 롤아웃은 훈련 지평을 훨씬 넘어서도 안정적으로 유지된다: 분포 품질은 우리가 측정한 가장 긴 지평인 5분까지 안정적으로 유지되며, 실제로 우리는 롤아웃이 붕괴 징후 없이 몇 시간 동안 지속되는 것을 관찰한다. 우리는 비디오 코덱, 생성 목적 함수, 멀티플레이어 조건화 방식이라는 핵심 설계 선택을 체계적으로 조사한다. 또한 모델 및 데이터 규모에 따라 행동이 어떻게 변화하는지, 즉 나타나는 능력과 지속되는 실패 모드를 특성화한다. 나아가 시각적 외관만이 아닌 모델의 물리적 이해를 평가하는 맞춤형 평가를 개발한다. 멀티플레이어 월드 모델에 대한 지속적인 연구를 지원하기 위해, 우리는 데이터셋, 전체 훈련 및 추론 코드베이스, 그리고 라이브 데모를 공개한다.
English
We introduce the first multiplayer world model for highly dynamic environments governed by complex physical interactions. Whereas single-player world models treat the other agents as part of the environment, ours conditions on the action streams of multiple agents, learning to attribute changes in the scene to the correct player and to stay coherent under arbitrary combinations of their actions. We study this problem in the game of Rocket League, where players compete and cooperate under fast, tightly coupled dynamics. Trained on 10,000 hours of gameplay collected with publicly available bots, our 5-billion-parameter latent diffusion model generates four-player matches in real time, producing 20 frames per second on a single Nvidia B200 GPU. Although trained only on short clips, its rollouts stay stable far beyond the training horizon: distributional quality holds steady out to five minutes, the longest horizon we measure, and in practice we observe rollouts continuing for hours with no sign of collapse. We systematically investigate the central design choices: the video codec, the generative objective, and the multiplayer conditioning scheme. In addition, we characterize how behavior changes with model and data scale, including the capabilities that emerge and the failure modes that persist. We further develop targeted evaluations that probe the model's physical understanding rather than visual appearance alone. To support continued research on multiplayer world models, we release our dataset, our full training and inference codebase, and a live demo.