ChatPaper.aiChatPaper

Business Arena:在逼真市場環境中評測 LLM 智能體

Business Arena: Benchmarking LLM Agents in a Realistic Marketplace

August 9, 2026
作者: Yijun Pan, Yukun Lian, Kunyu Shi, Junbo Li, Hongwei Xue, Sicong Xie, Guannan Zhang, Xiaoying Xing
cs.AI

摘要

經營企業是一種具挑戰性的智能工作。經營者必須從不完整的信號中推斷機會,在不確定性下投入資本,適應瞬息萬變的市場中延遲顯現的結果,並在合法交易前滿足監管要求。前沿大型語言模型智能體日益能夠完成複雜的工作流程,但現有的智能體基準測試卻很少評估與商業相關的能力。我們引入 Business Arena,這是一個受控環境,讓 AI 智能體經營一家跨境商店,在長期的時間跨度內向供應商採購並向買家銷售。我們以真實的阿里巴巴採購數據以及根據權威來源校準的市場條件作為此環境的基礎。延遲且相互關聯的後果使單一商業決策難以評判,但其綜合結果可透過利潤來衡量。由於僅靠利潤無法解釋智能體成功或失敗的原因,我們將智能體與人類設計的策略進行比較以估算可用機會,使用技能層級指標來揭示潛在的優勢與劣勢,並將已實現的收益與損失追溯到產生它們的行動。我們使用機制消融來確立,出色的結果反映的是真正的商業智慧,而非投機取巧或模擬器特有的捷徑。我們評估了 15 個前沿模型,發現平均最終淨資產存在九倍差異。即使是最優秀的模型也落後於人類設計的策略,這表明商業經營對大型語言模型智能體而言仍具挑戰性。技能層級分析揭示了不同的經營風格,從注重利潤率的溢價賣家到高周轉的批發商和客戶服務專家,而行動層級的歸因則識別出創造或摧毀價值的採購、定價和補救決策。總而言之,Business Arena 為評估端到端商業智能體邁出了通往真實且可信測試環境的第一步。
English
Running a business is a challenging form of intelligent work. Operators must infer opportunities from partial signals, commit capital under uncertainty, adapt to delayed outcomes in a changing market, and satisfy regulatory obligations before trading legally. Frontier LLM agents can increasingly complete complex workflows, yet business-related capabilities are rarely evaluated in existing agent benchmarks. We introduce Business Arena, a controlled environment where an AI agent runs a cross-border shop, buying from suppliers and selling to buyers over a long horizon. We ground the arena in real Alibaba.com sourcing data and market conditions calibrated from authoritative sources. Delayed and coupled consequences make individual business decisions difficult to judge, but their combined outcome is measurable through profit. Because profit alone cannot explain why an agent succeeds or fails, we compare agents with human-designed strategies to estimate available opportunity, use skill-level metrics to reveal underlying strengths and weaknesses, and trace realized gains and losses to the actions that produced them. We use mechanism ablations to establish that strong results reflect genuine business intelligence rather than neglect or simulator-specific shortcuts. We evaluate 15 frontier models and find a ninefold difference in mean final net worth. Even the best model falls behind human-designed strategies, indicating that business operation remains challenging for LLM agents. Skill-level analysis reveals operating styles, from margin-focused premium sellers to high-turnover wholesalers and customer-service specialists, while action-level attribution identifies the sourcing, pricing, and recovery decisions that create or destroy value. Together, Business Arena takes a first step toward a realistic and trustworthy testbed for evaluating end-to-end business agents.