ChatPaper.aiChatPaper

FM-Bench: 경쟁 에이전트를 포함한 장기 지평 관리 벤치마크

FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents

August 19, 2026
저자: Tianyou Wang, Chongyang Gao, Kezhen Chen, Chen Dong, Yinghao He, Donghan Li, Wangcheng Xu, Hongjiu Zhang, Chi Li
cs.AI

초록

언어 모델 에이전트는 이제 제한된 작업을 안정적으로 수행한다. 그러나 행동이 누적적 결과를 낳고 환경이 그 선택에 반응하는 긴 시간 지평에서 효과적인 의사 결정을 유지할 수 있는지는 대체로 측정되지 않은 채 남아 있다. FM-Bench(축구 경영 벤치마크)는 이를 측정한다. LLM 에이전트는 26개의 도구와 약 340~400개의 의사 결정 지점을 거쳐 게임 내 20년 동안 축구 클럽을 운영한다. 모든 경쟁 팀과 동일한 예산으로 선수단을 구성하고, 선수를 트레이드하며, 계약을 협상하고, 시설과 유소년 육성에 투자하고, 라인업을 짜며, 자신을 해고할 수 있는 이사회에 보고한다. 그동안 결정론적 엔진은 매년의 결과를 최종 점수 하나로 누적하며, LLM 심사자나 인간 평가자는 개입하지 않는다. 단독 트랙에서는 15개의 최첨단 모델 각각을 고정된 스크립트 세계와 대결시키고, 아레나에서는 동일한 모델들에 스크립트 기반 기준 모델을 더해 하나의 공유된 20년 세계에 배치한다. 우리가 아는 한, 이 규모의 최초 맞대결 평가다. 우리는 점수를 뒷받침하는 여섯 가지 행동 역량을 측정한다. 세 개의 시드에 걸쳐 15개 모델 모두 모든 지평을 완료한 반면, 스크립트 기반 블라인드 기준선은 대부분의 지평에서 소멸했다. claude-fable-5는 평균 점수에서 단독 트랙 보드 1위를 차지했고 아레나에서도 1위를 했지만, 아레나의 우승 타이틀은 그럼에도 10개 모델 사이를 순환했다. 규모, 가격, 공급업체 중 어느 것도 순위를 예측하지 못한다. 순위는 지평 후반에 가서야 확정되며, 첫 플레이에서 최고 성적을 낸 인간은 모델 순위표의 최하위에 그쳤다. 모델들을 구분 짓는 것은 계산 능력이 아니라 관리 행동이다. 고득점 모델은 끝에 가까울수록 수익이 더딘 투자를 줄이고, 현금을 유휴 상태로 두지 않고 투자 상태로 유지하며, 마감 훨씬 전에 재계약을 시작한다. 반면 토큰 사용량은 아무것도 예측하지 못한다. 어떤 모델도 수백 건의 거절된 입찰에서 시장의 숨겨진 가격을 학습하지 못하며, 자체 관리 메모리는 서로 반대되는 두 가지 방식으로 실패한다. 계속 증가만 하는 아카이브와 매 시즌 다시 쓰이는 계획이 그것이다. 코드는 https://github.com/Analogy-AI/fm-bench에서 이용할 수 있다.
English
Language model agents now execute bounded tasks reliably. Whether they can sustain effective decision-making over long horizons, where actions have cumulative consequences and the environment responds to their choices, remains largely unmeasured. FM-Bench (Football Management Benchmark) measures this. An LLM agent runs a football club for 20 in-game years through 26 tools and roughly 340 to 400 decision stops. It drafts a squad on the same budget as every rival, trades players, negotiates contracts, invests in facilities and youth, sets lineups, and answers to a board that can fire it, while a deterministic engine accumulates every year into one final score with no LLM judge or human rater. The solo track plays each of 15 frontier models against a frozen scripted world, and the Arena places the same models plus a scripted anchor in one shared 20-year world; to our knowledge, the first head-to-head evaluation at this scale. We measure six behavioral capabilities behind the score. Across three seeds, all 15 models complete every horizon while the blind scripted baselines die out in most of theirs, and claude-fable-5 tops the solo board on mean score and the Arena, where the title nonetheless rotates among ten models. Neither scale, price, nor vendor predicts the order; the order settles only late in the horizon, and the best first-play human lands only at the bottom of the model board. What separates the models is managerial behavior rather than computation. Higher-scoring models reduce slow-payoff investment near the end, keep cash invested rather than idle, and open renewals well before the deadline, while token spend predicts nothing. No model learns the market's hidden prices from hundreds of rejected bids, and self-managed memory fails in two opposite modes: an archive that only grows or a plan rewritten every season. Code is available at https://github.com/Analogy-AI/fm-bench.