Researchers released E-Commerce Bench, an open-source test of AI agents running online stores. The agent manages several shops across a simulated 365-day year. Tasks include market research, supplier negotiation, order fulfilment, returns and cash flow. Product and supplier data come from a real e-commerce platform. A calendar of promotions, natural disasters and supply-chain shocks keeps changing demand. Customer demand and supplier bargaining follow fixed rules, so runs can be repeated. The team scored 18 frontier models on seven measures and found no overall winner. GPT-5.6 Sol earned the most money but placed 16th of 18 on avoiding fraud. Among openly downloadable models, Qwen3.8-Max-Preview led and improved its bargaining over repeated orders. The code is published on GitHub.
What changed
Existing agent benchmarks chained short tasks over more turns, without year-long dynamics.
What it unlocks
Testing whether an AI agent can run a business over thousands of steps, not one task.
- 18 models evaluated
- 365-day simulated year
- GPT-5.6 Sol: 100,000 to 1,431,425
- Qwen3.8-Max-Preview: 416,252
Sources