ChatPaper.aiChatPaper

Business Arena: 현실적 시장 환경에서의 LLM 에이전트 벤치마킹

Business Arena: Benchmarking LLM Agents in a Realistic Marketplace

August 9, 2026
저자: Yijun Pan, Yukun Lian, Kunyu Shi, Junbo Li, Hongwei Xue, Sicong Xie, Guannan Zhang, Xiaoying Xing
cs.AI

초록

사업을 운영하는 것은 난이도가 높은 지적 작업이다. 운영자는 불완전한 신호에서 기회를 추론하고, 불확실성 속에서 자본을 투입하며, 변화하는 시장에서 지연되어 나타나는 결과에 적응하고, 합법적으로 거래하기 전에 규제 요건을 충족해야 한다. 최첨단 LLM 에이전트는 점점 더 복잡한 워크플로우를 완수할 수 있지만, 비즈니스 관련 역량은 기존 에이전트 벤치마크에서 거의 평가되지 않는다. 우리는 AI 에이전트가 국경 간 상점을 운영하며 장기간에 걸쳐 공급업체로부터 구매하고 구매자에게 판매하는 통제된 환경인 Business Arena를 소개한다. 우리는 이 아레나를 실제 Alibaba.com 소싱 데이터와 권위 있는 출처에서 보정한 시장 조건에 기반하여 구축했다. 지연되고 상호 연계된 결과는 개별 비즈니스 결정을 판단하기 어렵게 만들지만, 결합된 결과는 이윤을 통해 측정할 수 있다. 이윤만으로는 에이전트의 성공과 실패 이유를 설명할 수 없으므로, 우리는 에이전트를 인간이 설계한 전략과 비교하여 가용 기회를 추정하고, 스킬 수준 지표를 사용해 근본적인 강점과 약점을 드러내며, 실현된 이익과 손실을 이를 만들어낸 행동까지 추적한다. 우리는 메커니즘 제거 실험을 통해 강력한 결과가 과제를 무시하거나 시뮬레이터 특유의 지름길을 이용한 데서 비롯된 것이 아니라 진정한 비즈니스 지능을 반영한다는 것을 입증한다. 우리는 15개의 최첨단 모델을 평가했고, 평균 최종 순자산에서 9배 차이를 발견했다. 최고의 모델조차 인간이 설계한 전략에 뒤처지며, 이는 LLM 에이전트에게 비즈니스 운영이 여전히 어려운 과제임을 시사한다. 스킬 수준 분석은 마진 중심의 프리미엄 판매자부터 높은 회전율의 도매상, 고객 서비스 전문가에 이르는 운영 스타일을 드러내고, 행동 수준의 기여도 분석은 가치를 창출하거나 파괴하는 소싱, 가격 책정, 회복 결정을 식별한다. 이로써 Business Arena는 엔드투엔드 비즈니스 에이전트를 평가하기 위한 현실적이고 신뢰할 수 있는 테스트베드를 향한 첫 걸음이 된다.
English
Running a business is a challenging form of intelligent work. Operators must infer opportunities from partial signals, commit capital under uncertainty, adapt to delayed outcomes in a changing market, and satisfy regulatory obligations before trading legally. Frontier LLM agents can increasingly complete complex workflows, yet business-related capabilities are rarely evaluated in existing agent benchmarks. We introduce Business Arena, a controlled environment where an AI agent runs a cross-border shop, buying from suppliers and selling to buyers over a long horizon. We ground the arena in real Alibaba.com sourcing data and market conditions calibrated from authoritative sources. Delayed and coupled consequences make individual business decisions difficult to judge, but their combined outcome is measurable through profit. Because profit alone cannot explain why an agent succeeds or fails, we compare agents with human-designed strategies to estimate available opportunity, use skill-level metrics to reveal underlying strengths and weaknesses, and trace realized gains and losses to the actions that produced them. We use mechanism ablations to establish that strong results reflect genuine business intelligence rather than neglect or simulator-specific shortcuts. We evaluate 15 frontier models and find a ninefold difference in mean final net worth. Even the best model falls behind human-designed strategies, indicating that business operation remains challenging for LLM agents. Skill-level analysis reveals operating styles, from margin-focused premium sellers to high-turnover wholesalers and customer-service specialists, while action-level attribution identifies the sourcing, pricing, and recovery decisions that create or destroy value. Together, Business Arena takes a first step toward a realistic and trustworthy testbed for evaluating end-to-end business agents.