Research2026-09-06

ARC Prize scored OpenAI's GPT-6 Astra two ways on its ARC-AGI-3 test. Its own standard software gave 62.7%. OpenAI's Provider Adapter gave 99.9%, and cost less to run. The adapter keeps the model's reasoning state between requests and compresses long conversations. With the reasoning dial set to none, the adapter still returned 96.7%. ARC Prize said it is not claiming this is general intelligence. It will publish both harness results side by side in future. Fortune compared saved copies of OpenAI's launch post and found five figures changed after publication. OpenAI said most evaluations carry a few points of noise. Artificial Analysis ran separate tests and found Astra level with its predecessor on general intelligence, at higher cost.

What changed

OpenAI's 99.9% ARC-AGI-3 score was cited as evidence of human-level general ability.

What it unlocks

Comparing model scores like-for-like by asking which harness produced them.

  • 62.7% standard harness vs 99.9% adapter
  • 96.7% at zero reasoning effort
  • $26,098 vs $18,817 per run
  • five metrics revised after launch

Send this to someone who needs it

Shares the story and its sources. Nothing about you.

What does this mean for your job?

This is the story as everyone gets it. Once a week we send you the version written for your role — what changed, why it matters for the work you actually do, and one thing to try. Free while we tune it.