HarnessEval-W: 시각적 세계 평가의 에이전트화
HarnessEval-W: Agentifying the Evaluation of Visual Worlds
August 17, 2026
저자: Weiliang Chen, Haowen Sun, Jun Gao, Jiawei Chi, Hanyang Wang, Qiyu Dai, Yihao Li, Hao Li, Jingnan Gao, Yi-Hsin Hung, Xingzhuo Guo, Shangchen Miao, Zhiyuan Shi, Xiang Li, Fengrui Tian, Weihua Du, Ziqi Huang, Shenyuan Gao, Siqiao Huang, Mingyu Liu, Yifei Li, Shizun Wang, Xi Wang, Tianqi Zhang, Xue Luo, Xiyin Ren, Jinshan Ren, Xiaoyang Shen, Xiaobo Hu, Zhiyang Dou, Mingyu Ding, Yichao Yan, Xinchao Wang, Yizhou Wang, Shilong Liu, Wenzhao Zheng, Yueqi Duan, Yuan Gong, Ziwei Liu, Ming-Yu Liu, Jialong Wu, Jiangran Lyu, Fangfu Liu
cs.AI
초록
벤치마크는 단순한 스칼라 점수 이상을 제공해야 한다. 평가의 신뢰성을 보장하는 것은 점수를 정당화하는 추론이다. 이는 특히 세계 모델에서 중요한데, 롤아웃을 판단하려면 물리, 인과성, 세계 상태가 올바르게 진화하는지 이해해야 하기 때문이다. 인간은 이러한 위반을 자연스럽게 발견하지만, 기존의 어떤 벤치마크도 이 능력을 자동화하지 못한다. 메트릭은 무차별적으로 계산되며, 검토하거나 검증할 수 있는 추론 체인이 남지 않는다. 우리는 LLM 생태계의 하네스 패러다임을 세계 모델 벤치마킹에 도입한 에이전트화된 평가 파이프라인인 HarnessEval-W를 소개한다. HarnessEval-W는 고정된 루브릭을 적용하는 대신 각 평가 사례의 맥락을 해석하고, 평가 질문을 측정 가능한 하위 문제들로 분해하며, 각각의 하위 문제에 대해 추론할 수 있도록 맞춤화된 맥락과 진단 도구를 갖춘 전문화된 하위 에이전트들을 생성한다. 이후 상위 에이전트는 수집된 증거를 검증하고 이를 최종 판정으로 요약한다. 이러한 계층적 워크플로는 모든 평가를 투명한 증거 트리로 전환하며, 그 완전한 추론 체인이 결과를 정당화한다. 우리는 HarnessEval-W를 330개의 평가 사례에 걸쳐 18개의 대표적인 세계 모델에 적용했다. 이 평가는 생성된 모든 롤아웃에 대해 검증 가능하고 세밀한 진단을 제공하면서 인간의 선호와 밀접하게 일치한다. 우리는 전체 파이프라인을 라이브 벤치마크로 오픈소스화하며, 세계 모델이 진화함에 따라 새로운 기술과 평가 사례를 확장하는 데 기여할 광범위한 커뮤니티의 참여를 환영한다.
English
A benchmark should deliver more than a scalar score: what makes an evaluation trustworthy is the reasoning that justifies the score. This is especially critical for world models, where judging a rollout requires understanding whether physics, causality, and world state evolve correctly. Humans spot such violations naturally, yet no existing benchmark automates this capability: metrics are computed brute-force, leaving no reasoning chain that can be examined or verified. We introduce HarnessEval-W, an agentified evaluation pipeline that brings the harness paradigm from the LLM ecosystem to world model benchmarking. Rather than applying a fixed rubric, HarnessEval-W interprets the context of each evaluation case, decomposes the evaluation question into measurable subproblems, and spawns specialized sub-agents, each equipped with tailored context and diagnostic tools to reason over its own subproblem. The parent agent then validates the gathered evidence and summarizes it into the final verdict. This hierarchical workflow turns every evaluation into a transparent evidence tree whose complete reasoning chain justifies the result. We apply HarnessEval-W to 18 representative world models over 330 evaluation cases. Its judgments closely align with human preferences while providing verifiable, fine-grained diagnoses of every generated rollout. We open-source the full pipeline as a live benchmark and invite the broad community to contribute to grow new skills and evaluation cases as world models evolve.