Research2026-09-03

OpenAI reports that its GPT-6 Astra model scored 98.6% on ARC-AGI-3. That test is designed to measure general reasoning rather than memorised knowledge. The same benchmark returned 7.8% six months earlier. OpenAI has not disclosed the settings used during the test run. The model's reasoning steps are also not visible to outsiders. Those gaps make the result hard to verify independently. They also leave open what the score says about machine intelligence more broadly. The score and the caveats come from reporting by The New Stack.

What changed

The best score on this reasoning test six months ago was 7.8%.

  • 98.6% on ARC-AGI-3
  • 7.8% six months earlier

Send this to someone who needs it

Shares the story and its sources. Nothing about you.

What does this mean for your job?

This is the story as everyone gets it. Once a week we send you the version written for your role — what changed, why it matters for the work you actually do, and one thing to try. Free while we tune it.