MerchantBench: Eコマース業務におけるLLMエージェントの長期的な一貫性のベンチマーキング
MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations
July 31, 2026
著者: Qiming Shi, Yulong Tao, Linbo Jin, Zhaolu Kang, Yibo Dou, Jiawen Zhu, Tianjun Pan, Shaokang Fu, Chengyu Wang, Siyue Li, Yaping Cheng, Di Weng, Chengfu Huo
cs.AI
要旨
大規模言語モデルのエージェントは、自律的なツール利用者として評価されることがますます増えつつあるが、既存のベンチマークの大半は、成功基準が即時的な限定タスクに焦点を合わせている。実世界での応用には、長期的な一貫性(Long-Term Coherence)、すなわち長い期間にわたって目的を持った行動を維持し、蓄積された証拠に応じて意思決定を適応させる能力がしばしば必要とされる。この能力を評価するには、行動が将来の選択肢を制約し、フィードバックが不均一な遅延で届き、一貫性のない行動が測定可能な累積効果を生む持続的な環境が不可欠である。売り手側の電子商取引は、商品調達、出品・価格管理、キャッシュフロー管理、混合遅延フィードバック適応といった反復的かつ相互依存的な意思決定を通じて、この評価に適した環境を提供する。本稿では、98,843件の実在する電子商取引の商品レコードに基づき、エージェントとのインタラクションに26のツールを備えた365日間の注文レベルシミュレーションであるMerchantBenchを紹介する。MerchantBenchは、即時に観測可能な上流サプライヤーイベントと遅延のある下流の注文結果を結合し、エージェントに個々の注文ライフサイクルを追跡させ、過去の決定を再考させる。我々は2つのエージェントフレームワークの下で8つのLLMを、それぞれ365日のシミュレーションに及ぶ48回の実行で評価した。結果は、最新のLLMと人間の参加者との間に大きな隔たりがあることを示しており、最良のLLM構成は、人間の参加者が達成した平均最終純資産の27.3%しか達成しなかった。
English
Large language model agents are increasingly evaluated as autonomous tool users, yet most benchmarks focus on bounded tasks with immediate success criteria. Real-world deployments often require Long-Term Coherence, the capacity to preserve purposeful behavior across extended horizons while adapting decisions to accumulated evidence. Evaluating this capacity requires a persistent environment in which actions constrain future choices, feedback arrives at heterogeneous delays, and incoherent behavior produces measurable cumulative effects. Seller-side e-commerce provides a suitable setting for this evaluation through recurrent and interdependent decisions over Product Sourcing, Listing and Pricing Control, Cash-Flow Management, and Mixed-Latency Feedback Adaptation. We introduce MerchantBench, a 365-day order-level simulation grounded in 98,843 real e-commerce product records and equipped with 26 tools for agent interaction. MerchantBench couples promptly observable Upstream Supplier Events with delayed Downstream Order Outcomes, requiring agents to follow individual order lifecycles and revisit earlier decisions. We evaluate eight LLMs under two agent frameworks in 48 runs, each spanning 365 simulated days. Our results reveal a substantial gap between even the latest LLMs and human participants, with the best LLM configuration attaining only 27.3\% of the mean final net assets achieved by human participants.