Skip to content
Hacker News front page

Epoch AI and METR launch MirrorCode to test AI on reimplementing full software projects

What's the largest software project AI can complete on its own?

MirrorCode is a new benchmark where AI must reimplement 25 full programs from scratch without seeing the source code, matching the original output exactly on end-to-end tests. Tasks span Unix utilities, data serialization, bioinformatics, interpreters, static analysis, cryptography, and compression. Unlike existing benchmarks, it provides a real inference budget: the most expensive run cost $2,600 and the AI worked for 19 days without human intervention. Epoch AI estimates a human engineer without AI would need months for the hardest tasks. The benchmark is cheat-resistant by design, though the post doesn't detail the mechanism.

Why it matters: Epoch AI's MirrorCode benchmark measures AI's ability to independently ship complete software projects via end-to-end test parity, with real inference budgets instead of fixed token caps. Covers 25 tasks across 7 categories, with one run costing $2,600 over 19 days, and includ...

Read the original ↗Export Markdown