Research2026-08-31

Researchers released E-Commerce Bench, an open-source test of AI agents running online stores. The agent manages several shops across a simulated 365-day year. Tasks include market research, supplier negotiation, order fulfilment, returns and cash flow. Product and supplier data come from a real e-commerce platform. A calendar of promotions, natural disasters and supply-chain shocks keeps changing demand. Customer demand and supplier bargaining follow fixed rules, so runs can be repeated. The team scored 18 frontier models on seven measures and found no overall winner. GPT-5.6 Sol earned the most money but placed 16th of 18 on avoiding fraud. Among openly downloadable models, Qwen3.8-Max-Preview led and improved its bargaining over repeated orders. The code is published on GitHub.

What changed

Existing agent benchmarks chained short tasks over more turns, without year-long dynamics.

What it unlocks

Testing whether an AI agent can run a business over thousands of steps, not one task.

  • 18 models evaluated
  • 365-day simulated year
  • GPT-5.6 Sol: 100,000 to 1,431,425
  • Qwen3.8-Max-Preview: 416,252

Sources

Send this to someone who needs it

Shares the story and its sources. Nothing about you.

What does this mean for your job?

This is the story as everyone gets it. Once a week we send you the version written for your role — what changed, why it matters for the work you actually do, and one thing to try. Free while we tune it.