Research2026-08-03

Researchers at Epoch Research, led by Tom Adamczewski with six co-authors, released MirrorCode, a long-horizon coding benchmark in which AI agents must reimplement entire software projects from observed behavior alone, without access to the original source code. Solutions must match the original program's output exactly on end-to-end tests, including held-out tests. The benchmark covers 25 target programs across Unix utilities, data serialization and query tools, bioinformatics, interpreters, static analysis, cryptography and compression. The strongest model scored 56% across the benchmark, and agents reimplemented gotree, a 16,000-line bioinformatics toolkit the authors estimate would take a human engineer weeks. Evaluating the frontier required unusually large inference budgets, including $2,600 over 19 days for a single attempt on one large task. The paper was submitted 29 June 2026 and revised 17 July 2026, with code on GitHub.

Send this to someone who needs it

Shares the story and its sources. Nothing about you.

What does this mean for your job?

This is the story as everyone gets it. Once a week we send you the version written for your role — what changed, why it matters for the work you actually do, and one thing to try. Free while we tune it.