WorldCupArena: 언어 모델 및 심층 연구 에이전트의 축구 예측에 대한 세밀한 평가
WorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Football Forecasting
July 20, 2026
저자: Zhaokai Wang, Tianlin Gui, Jiayuan Rao, Shangzhe Di, Yihong Tang, Dingli Liang
cs.AI
초록
경기 시작 전에 축구 경기를 예측하려면 과거 결과만 아는 것 이상이 필요하다. 모델은 변화하는 정보를 활용하고 정답이 밝혀지기 전에 명확한 예측을 내놓아야 한다. 우리는 언어 모델과 심층 연구 에이전트를 위한 동적 벤치마크인 WorldCupArena를 소개한다. 2026년 FIFA 월드컵이 첫 번째 평가 대상이며, 동일한 절차는 향후 리그와 컵 대회에도 재사용될 수 있다. 각 경기 전에 모델은 공통 증거 패키지를 받거나 스스로 정보를 검색하여 결과와 점수, 예상 선수 및 이벤트, 경기 통계, 대회 결과를 예측한다. 경기 후에는 이 예측들이 기록된 결과와 비교된다. 우리는 결과 정확도, 정확한 점수 정확도, 예측 점수가 정확하지는 않지만 근접할 때 일부 점수를 부여하는 스코어라인 점수(Scoreline Score)를 보고하며, 다른 예측 과제에 대한 점수도 함께 제시한다. 104경기와 13개 시스템에 걸쳐, 유사한 결과 정확도를 가진 모델들도 세부 예측에서는 더 명확한 차이를 보였다. 베팅 시장 및 일반 팬 기준선과 비교했을 때, 최고 시스템은 결과 정확도와 정확한 점수 정확도에서는 작은 향상에 그쳤지만, 스코어라인에서는 더 뚜렷한 향상을 보였다. 새로운 일정이 시작됨에 따라 추가할 수 있어, 이미 알려진 결과를 사용하지 않고도 미래 모델을 평가할 수 있다. 코드, 프롬프트, 예측 및 평가 스크립트는 https://github.com/wzk1015/WorldCupArena에서 오픈소스로 공개된다.
English
Predicting a football match before kickoff requires more than knowing past results: a model must use changing information and make a clear prediction before the answer is available. We present WorldCupArena, a dynamic benchmark for language models and deep-research agents. The 2026 FIFA World Cup is its first evaluation, and the same process can be reused for future leagues and cups. Before each match, a model either receives a common evidence package or searches for information itself. It predicts the result and score, likely players and events, match statistics, and the outcome of the competition. After the match, these predictions are compared with the recorded result. We report result accuracy, exact-score accuracy, and a scoreline score that gives some credit when a predicted score is close but not exact, together with scores for the other prediction tasks. Across 104 matches and 13 systems, models with similar result accuracy differ more clearly on detailed predictions. Compared with betting-market and human-fan baselines, the best system shows only small gains in result and exact-score accuracy, but a clearer gain in Scoreline. New schedules can be added as they begin, allowing the benchmark to evaluate future models without using outcomes that are already known. Code, prompts, predictions, and evaluation scripts are open sourced at https://github.com/wzk1015/WorldCupArena.