Epoch AI and METR launch MirrorCode to test AI on reimplementing full software projects
What's the largest software project AI can complete on its own?
MirrorCode is a new benchmark where AI must reimplement 25 full programs from scratch without seeing the source code, matching the original output exactly on end-to-end tests. Tasks span Unix utilities, data serialization, bioinformatics, interpreters, static analysis, cryptography, and compression. Unlike existing benchmarks, it provides a real inference budget: the most expensive run cost $2,600 and the AI worked for 19 days without human intervention. Epoch AI estimates a human engineer without AI would need months for the hardest tasks. The benchmark is cheat-resistant by design, though the post doesn't detail the mechanism.
Why it matters: Epoch AI's MirrorCode benchmark measures AI's ability to independently ship complete software projects via end-to-end test parity, with real inference budgets instead of fixed token caps. Covers 25 tasks across 7 categories, with one run costing $2,600 over 19 days, and includ...