ARC Prize published verified test results for Google's Gemini 3.6 Flash across its four reasoning effort settings on the ARC-AGI-1 and ARC-AGI-2 abstract puzzle benchmarks. At the highest effort the model solved 91.2% of ARC-AGI-1 tasks and 60.4% of ARC-AGI-2 tasks, with per-task costs listed for each. Scores fall sharply at lower effort settings, dropping to 2.6% on ARC-AGI-2 at the minimal setting, and no ARC-AGI-3 results were reported.
What it unlocks
Comparing accuracy and cost per task across four reasoning effort settings before choosing one for a workload.
- 91.2% on ARC-AGI-1 Semi-Private at high effort
- 60.4% on ARC-AGI-2 Semi-Private at high effort
- $0.34 per task on ARC-AGI-1, $0.61 per task on ARC-AGI-2
- Minimal reasoning: 34.5% and 2.6%
- arcprize.org2026-08-06