An essay on Latent.Space argues that the software wrapped around AI models — the tools, memory and permission rules that let a model act — is steadily being trained into the models themselves, letting engineers delete much of it. It cites benchmark work where the same model scored between 52.4 and 76.2 depending only on the surrounding software, and Anthropic removing most of Claude Code's system instructions. The author predicts the remaining layer will govern when an agent may interrupt a person.
What changed
Agent scaffolding was treated as external prompting and tooling bolted onto a fixed model.
What it unlocks
A framing for deciding which agent scaffolding to remove as models absorb it, and where to invest instead.
- 52.4 to 76.2 across harnesses
- 106 tasks, same model
- 13.3% to 38.3% on ARC-AGI-3
- 80% of Claude Code system prompt deleted
Sources