Artificial Analysis updated its Intelligence Index to version 4.2. The firm added AA-Briefcase, its own test of agent-style knowledge work built by industry experts. It also added GDP.pdf, a document reasoning test created by Surge AI over 100 long PDFs. GPQA Diamond was removed because top models now score near the maximum. Private test sets that labs cannot see now carry a larger share of the score. Grading was also reworked across several tests to reduce scoring errors. Anthropic's Claude Fable 5.1 leads the Index, ahead of OpenAI's GPT-6 Astra. Meta ranks third among labs, followed by SpaceXAI, Moonshot, Z.AI and Google. Artificial Analysis calls this an interim step before a larger version 5.
What changed
v4.1 included GPQA Diamond and kept 20% of weighting in private test sets.
What it unlocks
Comparing models on multi-week knowledge work projects and long professional document reasoning.
- 40% of weighting now held-out, double v4.1
- GDP.pdf spans 4,592 PDF pages
- GPT-6 Astra 33.2% on GDP.pdf
- GPT-6 Astra +4pts over GPT-5.6 Sol
Sources