Artificial Analysis published the methodology behind version 4.2 of its Intelligence Index. The index combines ten evaluations into a single score for language models. Weighting favours agent tasks at 30 percent, then general ability, coding and scientific reasoning. AA-Briefcase, a new test, sets multi-week knowledge work projects with thousands of source files. Models run in isolated sandboxes with no internet access. Three separate judge models grade rubric checks and head-to-head comparisons. Artificial Analysis also runs extra tests for multilingual, visual and legal work outside the index. It states the suite is mainly text-based and English-language.
What changed
Earlier versions of the index used a different set and weighting of evaluations.
What it unlocks
Comparing language models on a documented, weighted suite of agentic, coding and reasoning tests.
- 10 evaluations in Index v4.2
- Agents weighted 30% of index
- 95% confidence interval under ±1%
Sources