ChatPaper.aiChatPaper

PlayWorld: 장기적 목표에 대한 에이전트 플레이어 기반 세계 모델 벤치마킹

PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives

August 13, 2026
저자: Kaixin Ding, Xi Chen, Minghong Cai, Zhiyuan Xu, Yiyang Wang, Yuxiang Lu, Junyi Li, Shuyang Chen, Yuan Gao, Xin Tao, Pengfei Wan, Hengshuang Zhao
cs.AI

초록

비디오 세계 모델은 현재 관찰과 사용자 행동에 조건화된 미래 상태를 시뮬레이션한다. 최근 시스템들은 긴 시퀀스에 걸쳐 인상적인 비디오 일관성과 행동 제어 가능성을 입증했다. 그러나 이러한 상호작용 모델들을 공정하게 비교하는 것은 여전히 어려운 과제로 남아 있다. 실제로 인간 플레이어는 장기 목표를 추구하기 위해 상호작용함으로써 세계 모델을 평가한다. 예를 들어, 사용자는 환경이 일관성을 유지하는지 확인하기 위해 360도 회전하거나, 물속으로 걸어 들어가 현실적인 물결이 생성되는지 관찰할 수 있다. 동일한 목표를 달성하는 데 필요한 행동 시퀀스는 모델 간에 상당히 다를 수 있으므로, 고정된 행동 조건부 평가는 모델 간 비교에 적합하지 않다. 이를 해결하기 위해 우리는 멀티모달 에이전트 플레이어를 활용하여 지정된 장기 목표를 향해 세계 모델과 상호작용하게 한다. 이 패러다임을 기반으로 우리는 171개의 시나리오를 제공하는 벤치마크인 PlayWorld를 소개하며, 각 시나리오에는 명시된 목표가 포함된다. 성능을 철저히 평가하기 위해 우리는 기하학적 일관성, 상호작용 충실도, 시야 밖 진화, 통찰 진화라는 네 가지 핵심 차원을 따라 모델을 평가한다. 또한 비디오 품질과 제어 가능성을 위한 기본 능력 지표도 포함한다. 아홉 개의 최첨단 세계 모델에 대한 실험은 현재 모델들이 특히 공간 일관성 유지와 지속적인 상태 진화 측면에서 장기 상호작용 목표에 대해 여전히 불안정함을 보여준다. 코드와 데이터는 https://github.com/kxding/PlayWorld에서 확인할 수 있다.
English
Video world models simulate future states conditioned on current observations and user actions. Recent systems have demonstrated impressive video consistency and action controllability over long sequences. However, fairly comparing these interactive models remains challenging. In practice, a human player typically evaluates a world model by pursuing long-horizon objectives through interaction. For example, a user may turn around 360 degrees to see whether the environment remains consistent, or walk into the water and inspect whether realistic water ripples are generated. The action sequence required to achieve the same objective may vary substantially between models, making fixed action-conditioned evaluation unsuitable for cross-model comparison. To address this, we employ multi-modal Agent Players to interact with world models toward specified long-horizon objectives. Building on this paradigm, we introduce PlayWorld, a benchmark providing 171 scenarios, each with a specified objective. To evaluate performance thoroughly, we assess models along four core dimensions: geometry consistency, interaction fidelity, out-of-sight evolution, and insight evolution. In addition, we incorporate basic ability metrics for video quality and controllability. Experiments across nine state-of-the-art world models reveal that current models remain unreliable on long-horizon interactive objectives, particularly in maintaining spatial consistency and persistent state evolution. Code and data are available at https://github.com/kxding/PlayWorld.