MerchantBench: 전자상거래 운영에서 LLM 에이전트의 장기적 일관성 벤치마킹
MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations
July 31, 2026
저자: Qiming Shi, Yulong Tao, Linbo Jin, Zhaolu Kang, Yibo Dou, Jiawen Zhu, Tianjun Pan, Shaokang Fu, Chengyu Wang, Siyue Li, Yaping Cheng, Di Weng, Chengfu Huo
cs.AI
초록
대형 언어 모델 에이전트는 점점 더 자율적인 도구 사용자로서 평가되고 있지만, 대부분의 벤치마크는 즉각적인 성공 기준을 가진 제한된 작업에 초점을 맞추고 있다. 실제 배포 환경에서는 종종 장기적 일관성(Long-Term Coherence), 즉 축적된 증거에 따라 결정을 적응시키면서도 장기간에 걸쳐 목적 있는 행동을 유지하는 능력이 요구된다. 이러한 능력을 평가하려면 행동이 미래 선택을 제약하고, 피드백이 이질적인 지연 시간을 두고 도착하며, 비일관적인 행동이 측정 가능한 누적 효과를 초래하는 지속적인 환경이 필요하다. 판매자 중심 전자상거래는 제품 소싱, 등록 및 가격 통제, 현금 흐름 관리, 혼합 지연 피드백 적응에 걸친 반복적이고 상호 의존적인 의사 결정을 통해 이러한 평가에 적합한 환경을 제공한다. 우리는 98,843개의 실제 전자상거래 제품 기록을 기반으로 하고 에이전트 상호작용을 위한 26가지 도구를 갖춘 365일 주문 수준 시뮬레이션인 MerchantBench를 소개한다. MerchantBench는 즉시 관찰 가능한 업스트림 공급업체 이벤트와 지연된 다운스트림 주문 결과를 결합하여, 에이전트가 개별 주문 수명 주기를 추적하고 이전 결정을 재검토하도록 요구한다. 우리는 두 가지 에이전트 프레임워크 하에서 48회의 실행에 걸쳐 8개의 LLM을 평가했으며, 각 실행은 365일의 시뮬레이션 기간을 포함한다. 실험 결과는 최신 LLM조차도 인간 참가자와 상당한 격차를 보임을 보여주며, 최고 성능의 LLM 구성은 인간 참가자들이 달성한 평균 최종 순자산의 27.3%에 그쳤다.
English
Large language model agents are increasingly evaluated as autonomous tool users, yet most benchmarks focus on bounded tasks with immediate success criteria. Real-world deployments often require Long-Term Coherence, the capacity to preserve purposeful behavior across extended horizons while adapting decisions to accumulated evidence. Evaluating this capacity requires a persistent environment in which actions constrain future choices, feedback arrives at heterogeneous delays, and incoherent behavior produces measurable cumulative effects. Seller-side e-commerce provides a suitable setting for this evaluation through recurrent and interdependent decisions over Product Sourcing, Listing and Pricing Control, Cash-Flow Management, and Mixed-Latency Feedback Adaptation. We introduce MerchantBench, a 365-day order-level simulation grounded in 98,843 real e-commerce product records and equipped with 26 tools for agent interaction. MerchantBench couples promptly observable Upstream Supplier Events with delayed Downstream Order Outcomes, requiring agents to follow individual order lifecycles and revisit earlier decisions. We evaluate eight LLMs under two agent frameworks in 48 runs, each spanning 365 simulated days. Our results reveal a substantial gap between even the latest LLMs and human participants, with the best LLM configuration attaining only 27.3\% of the mean final net assets achieved by human participants.