HarnessEval-W: 視覚世界評価のエージェント化
HarnessEval-W: Agentifying the Evaluation of Visual Worlds
August 17, 2026
著者: Weiliang Chen, Haowen Sun, Jun Gao, Jiawei Chi, Hanyang Wang, Qiyu Dai, Yihao Li, Hao Li, Jingnan Gao, Yi-Hsin Hung, Xingzhuo Guo, Shangchen Miao, Zhiyuan Shi, Xiang Li, Fengrui Tian, Weihua Du, Ziqi Huang, Shenyuan Gao, Siqiao Huang, Mingyu Liu, Yifei Li, Shizun Wang, Xi Wang, Tianqi Zhang, Xue Luo, Xiyin Ren, Jinshan Ren, Xiaoyang Shen, Xiaobo Hu, Zhiyang Dou, Mingyu Ding, Yichao Yan, Xinchao Wang, Yizhou Wang, Shilong Liu, Wenzhao Zheng, Yueqi Duan, Yuan Gong, Ziwei Liu, Ming-Yu Liu, Jialong Wu, Jiangran Lyu, Fangfu Liu
cs.AI
要旨
ベンチマークはスカラースコア以上のものを提供すべきである。評価を信頼できるものにするのは、そのスコアを正当化する推論である。これはワールドモデルにとって特に重要であり、ロールアウトを評価するには、物理、因果関係、ワールド状態が正しく進化しているかを理解することが必要となる。人間はこのような違反を自然に察知するが、既存のベンチマークはこの能力を自動化していない。指標は力任せに計算され、検証可能な推論チェーンが残らない。我々は、LLMエコシステムのハーネスパラダイムをワールドモデルのベンチマークに導入する、エージェント化された評価パイプラインであるHarnessEval-Wを提案する。HarnessEval-Wは固定のルーブリックを適用するのではなく、各評価ケースの文脈を解釈し、評価問題を測定可能な部分問題に分解し、それぞれに調整された文脈と診断ツールを備えた専門化されたサブエージェントを生成して、各サブ問題について推論させる。その後、親エージェントが収集された証拠を検証し、最終判定に要約する。この階層的ワークフローにより、すべての評価は透明なエビデンスツリーとなり、その完全な推論チェーンが結果を正当化する。我々はHarnessEval-Wを、330の評価ケースにわたる18の代表的なワールドモデルに適用した。その判定は人間の好みと密接に一致し、生成されたすべてのロールアウトについて検証可能で詳細な診断を提供する。我々は完全なパイプラインをライブベンチマークとしてオープンソース化し、ワールドモデルの進化に伴い、新しいスキルや評価ケースの拡充に広くコミュニティが貢献することを歓迎する。
English
A benchmark should deliver more than a scalar score: what makes an evaluation trustworthy is the reasoning that justifies the score. This is especially critical for world models, where judging a rollout requires understanding whether physics, causality, and world state evolve correctly. Humans spot such violations naturally, yet no existing benchmark automates this capability: metrics are computed brute-force, leaving no reasoning chain that can be examined or verified. We introduce HarnessEval-W, an agentified evaluation pipeline that brings the harness paradigm from the LLM ecosystem to world model benchmarking. Rather than applying a fixed rubric, HarnessEval-W interprets the context of each evaluation case, decomposes the evaluation question into measurable subproblems, and spawns specialized sub-agents, each equipped with tailored context and diagnostic tools to reason over its own subproblem. The parent agent then validates the gathered evidence and summarizes it into the final verdict. This hierarchical workflow turns every evaluation into a transparent evidence tree whose complete reasoning chain justifies the result. We apply HarnessEval-W to 18 representative world models over 330 evaluation cases. Its judgments closely align with human preferences while providing verifiable, fine-grained diagnoses of every generated rollout. We open-source the full pipeline as a live benchmark and invite the broad community to contribute to grow new skills and evaluation cases as world models evolve.