INT21 reported that its SwarmOS coordination layer raised OpenAI's GPT-5.6-Sol from 13.3% to a full score on the public set of the ARC-AGI-3 reasoning benchmark, and lifted a cheaper model from zero to 56%. The company says no benchmark-specific tuning was used. The results cover only the public 25-environment set, not the private competition sets used for official rankings.
What changed
The same model scored 13.3% on the benchmark's public set on its own, and NVIDIA's comparable full score used a Claude model.
What it unlocks
Evidence that a cheaper model with a stronger coordination layer can match results previously reached only with the most expensive models.
- 13.3% to 100% RHAE with Sol
- 0% to 56% RHAE with Luna
- 25 environments, 183 levels
- 6,731 environment actions
What you need to act on it
- access to INT21's SwarmOS platform
Sources