ChatPaper.aiChatPaper

ビジネスアリーナ:現実的な市場環境におけるLLMエージェントのベンチマーキング

Business Arena: Benchmarking LLM Agents in a Realistic Marketplace

August 9, 2026
著者: Yijun Pan, Yukun Lian, Kunyu Shi, Junbo Li, Hongwei Xue, Sicong Xie, Guannan Zhang, Xiaoying Xing
cs.AI

要旨

事業運営は、困難な知的作業の一形態である。経営者は、部分的なシグナルから機会を推測し、不確実性の下で資本を投入し、変化する市場における遅延した結果に適応し、合法的に取引する前に規制上の義務を満たさなければならない。最先端のLLMエージェントは複雑なワークフローをますます完了できるようになっているが、既存のエージェントベンチマークではビジネス関連の能力が評価されることはほとんどない。我々は、AIエージェントが長期間にわたって越境ショップを運営し、サプライヤーから仕入れ、バイヤーに販売する、制御された環境であるBusiness Arenaを紹介する。このアリーナは、実際のAlibaba.comの調達データと、信頼できる情報源から較正された市場状況に基づいている。遅延を伴い相互に連動する結果により、個々のビジネス上の意思決定の評価は難しいが、それらの複合的な成果は利益を通じて測定可能である。利益だけではエージェントが成功するか失敗するかの理由を説明できないため、我々はエージェントを人間が設計した戦略と比較して利用可能な機会を推定し、スキルレベルの指標を用いて根底にある長所と短所を明らかにし、実現した利益と損失をそれを生み出した行動に遡って追跡する。我々はメカニズムの除去実験を用いて、優れた結果が、怠慢やシミュレータ固有の近道ではなく、真のビジネス知能を反映していることを立証する。我々は15の最先端モデルを評価し、平均最終純資産に9倍の差があることを見いだした。最良のモデルでさえ人間が設計した戦略に及ばず、LLMエージェントにとって事業運営が依然として困難であることを示している。スキルレベルの分析により、利益率重視のプレミアム販売者から高回転の卸売業者、カスタマーサービス特化型まで、運営スタイルが明らかになる一方、行動レベルの属性分析は、価値を生み出すか破壊する調達、価格設定、回収の意思決定を特定する。総合すると、Business Arenaは、エンドツーエンドのビジネスエージェントを評価するための、現実的で信頼性の高いテストベッドへの第一歩となる。
English
Running a business is a challenging form of intelligent work. Operators must infer opportunities from partial signals, commit capital under uncertainty, adapt to delayed outcomes in a changing market, and satisfy regulatory obligations before trading legally. Frontier LLM agents can increasingly complete complex workflows, yet business-related capabilities are rarely evaluated in existing agent benchmarks. We introduce Business Arena, a controlled environment where an AI agent runs a cross-border shop, buying from suppliers and selling to buyers over a long horizon. We ground the arena in real Alibaba.com sourcing data and market conditions calibrated from authoritative sources. Delayed and coupled consequences make individual business decisions difficult to judge, but their combined outcome is measurable through profit. Because profit alone cannot explain why an agent succeeds or fails, we compare agents with human-designed strategies to estimate available opportunity, use skill-level metrics to reveal underlying strengths and weaknesses, and trace realized gains and losses to the actions that produced them. We use mechanism ablations to establish that strong results reflect genuine business intelligence rather than neglect or simulator-specific shortcuts. We evaluate 15 frontier models and find a ninefold difference in mean final net worth. Even the best model falls behind human-designed strategies, indicating that business operation remains challenging for LLM agents. Skill-level analysis reveals operating styles, from margin-focused premium sellers to high-turnover wholesalers and customer-service specialists, while action-level attribution identifies the sourcing, pricing, and recovery decisions that create or destroy value. Together, Business Arena takes a first step toward a realistic and trustworthy testbed for evaluating end-to-end business agents.