ChatPaper.aiChatPaper

AutoWorldModel-Bench:以狀態為中心的自動化世界模型研究基準

AutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model Research

July 20, 2026
作者: Marjan Moodi, Xuankang Zhu, Fernando De Mesentier Silva, Harold Chaput, Mohammad Reza Taesiri
cs.AI

摘要

世界建模是一個尚未定型的領域:架構、訓練目標與狀態表示之間存在複雜的交互作用,且沒有任何單一方案能橫跨各種環境佔據主導地位。這使其成為以自主研究人員角色運作的 AI 編碼代理的理想試驗場——在這種情境中,改進方向並非事先指定,這與當前代理基準測試中佔主導地位的按規格工程任務形成對比。我們推出 AutoModel-Bench,這是一個閉環基準測試,其中前沿編碼代理在固定計算預算下自主改進所提供的世界模型起點程式。該基準測試涵蓋八個遊戲環境,統一採用結構化狀態表示——從每個遊戲中提取的真實實體狀態,並透過共享的張量格式進行消費——如此可將動力學建模與感知分離,使每次運行僅需數分鐘即可完成疊代。在 64 場會話中,Codex-5.4 與 Claude Opus 4.6 在其中 63 場成功改進了起點程式;在 91% 的會話中,勝出的編輯屬於非平凡的科研型修改——新目標、新表示、新 rollout 程序或架構變更——而非超參數微調。我們的基準測試提供了一個情境,使前沿編碼代理得以在開放式研究而非按規格工程問題上接受評估。
English
World modeling is an unsettled field: architectures, training objectives, and state representations interact in complex ways, and no single recipe dominates across environments. This makes it an ideal testbed for AI coding agents acting as autonomous researchers--a setting in which the improvement direction is not specified in advance, unlike the engineering-to-spec tasks that dominate current agent benchmarks. We introduce AutoWorldModel-Bench, a closed-loop benchmark in which frontier coding agents autonomously improve a provided world-model starter under a fixed compute budget. The benchmark spans eight game environments under a unified structured-state representation--ground-truth entity state extracted from each game and consumed through a shared tensor format--which isolates dynamics modeling from perception and enables minutes-per-run iteration. Across 64 sessions, Codex-5.4 and Claude Opus 4.6 improve their starter on 63; in 91% of sessions the winning edit is a non-trivial research-style modification--a new objective, representation, rollout procedure, or architectural change--rather than a hyperparameter tweak. Our benchmark offers a setting in which frontier coding agents can be evaluated on open-ended research rather than engineering-to-spec problems.