E-Commerce Bench: Evaluating LLM Agents on Long-Horizon Autonomous Business Operation

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过E-Commerce Bench评估大语言模型在长期自主商业运营中的表现,涵盖市场研究、谈判等多方面,以最大化年末总资产。
📝 Abstract
Long-horizon agentic tasks go beyond chaining short tasks over more interaction turns. Their evolving dynamic environments and long-range dependencies require Large Language Models (LLMs) to continually explore, learn from experience, and adapt their policies over thousands of steps. We introduce E-Commerce Bench, the first open-source benchmark that integrates multi-round counterpart negotiation and dynamic events into a year-long business operation. Over a 365-day year, an LLM agent concurrently runs multiple online stores, researching the market, negotiating with suppliers to source inventory, optimizing sales strategies, fulfilling orders, handling returns, and managing cash flow to maximize its end-of-year total assets. To construct a realistic merchant-side operating environment, the product and supplier data are derived from a real e-commerce platform, while a year-long calendar of promotions, natural disasters, and supply-chain shocks continually reshapes demand. For reproducibility, both sides of the market are deterministic: customer purchases and returns follow a fixed demand model, while a negotiation kernel determines supplier pricing, concessions, and decisions, with an LLM used only to verbalize them. We evaluate 18 frontier models across seven dimensions, including year-end assets, and find that no single model dominates. GPT-5.6 Sol earns the most, growing the 100,000 opening stake into 1,431,425, yet it ranks 16th of 18 on fraud avoidance and trails Fable5 in operational efficiency. Among open-weight models, Qwen3.8-Max-Preview leads with 416,252, 38% above GLM 5.2 (high), and achieves the strongest learning over the horizon, progressively bargaining down prices across repeated orders. Our code is available at https://github.com/QwenLM/E-CommerceBench.
Problem

Research questions and friction points this paper is trying to address.

Long-horizon agentic tasks
Large Language Models (LLMs)
E-Commerce Bench
Dynamic environments
Autonomous business operation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Long-Horizon Autonomous Business Operation
Large Language Models (LLMs)
Multi-Round Negotiation
Dynamic Events
Deterministic Market Behavior
🔎 Similar Papers
2023-08-22Frontiers Comput. Sci.Citations: 866