OpenAI reports that its GPT-6 Astra model scored 98.6% on ARC-AGI-3. That test is designed to measure general reasoning rather than memorised knowledge. The same benchmark returned 7.8% six months earlier. OpenAI has not disclosed the settings used during the test run. The model's reasoning steps are also not visible to outsiders. Those gaps make the result hard to verify independently. They also leave open what the score says about machine intelligence more broadly. The score and the caveats come from reporting by The New Stack.
What changed
The best score on this reasoning test six months ago was 7.8%.
- 98.6% on ARC-AGI-3
- 7.8% six months earlier
Sources