ChatPaper.aiChatPaper

PlayWorld:以智能體玩家針對長期目標評測世界模型

PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives

August 13, 2026
作者: Kaixin Ding, Xi Chen, Minghong Cai, Zhiyuan Xu, Yiyang Wang, Yuxiang Lu, Junyi Li, Shuyang Chen, Yuan Gao, Xin Tao, Pengfei Wan, Hengshuang Zhao
cs.AI

摘要

視頻世界模型根據當前觀測和用戶動作模擬未來狀態。近期的系統已在長序列上展現出令人矚目的視頻一致性和動作可控性。然而,公平地比較這些互動式模型仍然充滿挑戰。在實際應用中,人類玩家通常通過互動追求長期目標來評估世界模型。例如,用戶可能旋轉360度以檢查環境是否保持一致,或走入水中觀察是否生成逼真的水波紋。要達成相同目標所需的動作序列在不同模型之間可能差異很大,這使得固定的動作條件評估不適合跨模型比較。為了解決這個問題,我們採用多模態智能體玩家(Agent Players)與世界模型互動,以達成指定的長期目標。在此範式基礎上,我們提出了PlayWorld,一個提供171個場景的基準測試,每個場景都有明確的目標。為了全面評估性能,我們沿四個核心維度評估模型:幾何一致性、互動真實度、視野外演化,以及洞察演化。此外,我們還納入了視頻質量和可控性的基本能力指標。對九個最先進世界模型的實驗結果表明,目前的模型在長期互動目標上仍不可靠,特別是在維持空間一致性和持續狀態演化方面。代碼和數據可在 https://github.com/kxding/PlayWorld 獲取。
English
Video world models simulate future states conditioned on current observations and user actions. Recent systems have demonstrated impressive video consistency and action controllability over long sequences. However, fairly comparing these interactive models remains challenging. In practice, a human player typically evaluates a world model by pursuing long-horizon objectives through interaction. For example, a user may turn around 360 degrees to see whether the environment remains consistent, or walk into the water and inspect whether realistic water ripples are generated. The action sequence required to achieve the same objective may vary substantially between models, making fixed action-conditioned evaluation unsuitable for cross-model comparison. To address this, we employ multi-modal Agent Players to interact with world models toward specified long-horizon objectives. Building on this paradigm, we introduce PlayWorld, a benchmark providing 171 scenarios, each with a specified objective. To evaluate performance thoroughly, we assess models along four core dimensions: geometry consistency, interaction fidelity, out-of-sight evolution, and insight evolution. In addition, we incorporate basic ability metrics for video quality and controllability. Experiments across nine state-of-the-art world models reveal that current models remain unreliable on long-horizon interactive objectives, particularly in maintaining spatial consistency and persistent state evolution. Code and data are available at https://github.com/kxding/PlayWorld.