Research2026-09-04

Artificial Analysis updated its Intelligence Index to version 4.2. The firm added AA-Briefcase, its own test of agent-style knowledge work built by industry experts. It also added GDP.pdf, a document reasoning test created by Surge AI over 100 long PDFs. GPQA Diamond was removed because top models now score near the maximum. Private test sets that labs cannot see now carry a larger share of the score. Grading was also reworked across several tests to reduce scoring errors. Anthropic's Claude Fable 5.1 leads the Index, ahead of OpenAI's GPT-6 Astra. Meta ranks third among labs, followed by SpaceXAI, Moonshot, Z.AI and Google. Artificial Analysis calls this an interim step before a larger version 5.

What changed

v4.1 included GPQA Diamond and kept 20% of weighting in private test sets.

What it unlocks

Comparing models on multi-week knowledge work projects and long professional document reasoning.

  • 40% of weighting now held-out, double v4.1
  • GDP.pdf spans 4,592 PDF pages
  • GPT-6 Astra 33.2% on GDP.pdf
  • GPT-6 Astra +4pts over GPT-5.6 Sol

Send this to someone who needs it

Shares the story and its sources. Nothing about you.

What does this mean for your job?

This is the story as everyone gets it. Once a week we send you the version written for your role — what changed, why it matters for the work you actually do, and one thing to try. Free while we tune it.