Independent testing firm Artificial Analysis reports that Anthropic's newly released Claude Opus 5 tops its benchmark for multi-step office work such as research reports, presentations and spreadsheets, at about 20% less cost per task than the previous leader. Gains come mainly from accuracy and analysis rather than the look of the output, where a rival GPT model still scores higher. The strongest settings take over 25 minutes per task.
What changed
Claude Fable 5 led the benchmark at $22.30 per task.
What it unlocks
Running long multi-file research, report and spreadsheet tasks at a lower cost per task than the previous best-scoring model.
- 1720 vs 1574 Elo (Fable 5)
- $17.79 vs $22.30 per task
- high effort: $10.41 per task
- 36.2 min per task at max effort
What you need to act on it
- API or product access to Claude Opus 5
- willingness to wait 25+ minutes per task
Sources