What's the largest software project AI can complete on its own?

Researchers have developed MirrorCode, a benchmark designed to test AI models on long-horizon software engineering tasks that require reimplementing entire programs. The study demonstrates that current models like Claude Opus can successfully complete complex, multi-week coding projects without human intervention.
Why it matters
This indicates a significant shift in AI capabilities from simple code snippets to autonomous, end-to-end software development.
AI has made rapid progress on software engineering benchmarks in the past few years. However, most such benchmarks tend to focus on shorter tasks like fixing bugs or implementing individual features. MirrorCode is our benchmark, co-developed with METR, to test AI models on long-horizon coding tasks. In a MirrorCode task, AI models are tasked with reimplementing an entire program end-to-end, without access to the original source code. AI-generated solutions must match the original program’s output exactly on end-to-end tests, including held-out tests. MirrorCode’s 25 target programs span different areas of computing: Unix utilities, data serialization and query tools, bioinformatics, interpreters, static analysis, cryptography, and compression.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in