MerchantBench:用于评估电商运营中LLM智能体长期连贯性的基准测试
MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations
July 31, 2026
作者: Qiming Shi, Yulong Tao, Linbo Jin, Zhaolu Kang, Yibo Dou, Jiawen Zhu, Tianjun Pan, Shaokang Fu, Chengyu Wang, Siyue Li, Yaping Cheng, Di Weng, Chengfu Huo
cs.AI
摘要
大语言模型智能体日益被评估为自主工具使用者,然而大多数基准测试聚焦于具有即时成功标准的受限任务。现实世界部署往往要求长期连贯性,即在扩展的时间跨度内保持有目的的行为,同时根据累积证据调整决策的能力。评估这一能力需要一个持久化环境,其中行动会约束未来的选择,反馈以异构延迟到达,且不连贯行为会产生可测量的累积效应。卖家侧电子商务通过商品选品、上架与定价控制、现金流管理以及混合延迟反馈适应中的循环性和相互依赖决策,为这种评估提供了合适的场景。我们提出MerchantBench,一个基于98,843条真实电子商务商品记录、配备26种工具供智能体交互的365天订单级模拟。MerchantBench将可即时观察的上游供应商事件与延迟的下游订单结果相耦合,要求智能体跟踪单个订单生命周期并重新审视早期决策。我们在两种智能体框架下对八种大语言模型进行了48次运行评估,每次运行跨越365个模拟日。结果表明,即使是最新的大语言模型与人类参与者之间仍存在显著差距,最佳大语言模型配置仅达到人类参与者平均最终净资产的27.3%。
English
Large language model agents are increasingly evaluated as autonomous tool users, yet most benchmarks focus on bounded tasks with immediate success criteria. Real-world deployments often require Long-Term Coherence, the capacity to preserve purposeful behavior across extended horizons while adapting decisions to accumulated evidence. Evaluating this capacity requires a persistent environment in which actions constrain future choices, feedback arrives at heterogeneous delays, and incoherent behavior produces measurable cumulative effects. Seller-side e-commerce provides a suitable setting for this evaluation through recurrent and interdependent decisions over Product Sourcing, Listing and Pricing Control, Cash-Flow Management, and Mixed-Latency Feedback Adaptation. We introduce MerchantBench, a 365-day order-level simulation grounded in 98,843 real e-commerce product records and equipped with 26 tools for agent interaction. MerchantBench couples promptly observable Upstream Supplier Events with delayed Downstream Order Outcomes, requiring agents to follow individual order lifecycles and revisit earlier decisions. We evaluate eight LLMs under two agent frameworks in 48 runs, each spanning 365 simulated days. Our results reveal a substantial gap between even the latest LLMs and human participants, with the best LLM configuration attaining only 27.3\% of the mean final net assets achieved by human participants.