AI has made accelerated advancement connected package engineering benchmarks successful the past fewer years. However, astir specified benchmarks thin to attraction connected shorter tasks for illustration fixing bugs aliases implementing individual features. MirrorCode is our benchmark, co-developed pinch METR, to trial AI models connected long-horizon coding tasks. In a MirrorCode task, AI models are tasked pinch reimplementing an full programme end-to-end, without entree to the original root code. AI-generated solutions must lucifer the original program’s output precisely connected end-to-end tests, including held-out tests. MirrorCode’s 25 target programs span different areas of computing: Unix utilities, information serialization and query tools, bioinformatics, interpreters, fixed analysis, cryptography, and compression.
How MirrorCode is different
Scale-aware evaluations
Crucially, we supply a ample capable conclusion fund to make a superior effort astatine MirrorCode tasks. Many existing package engineering benchmarks limit conclusion spending to astir $1–10, moreover erstwhile the task would return weeks for a quality to complete. For example, 1 of the largest MirrorCode tasks costs $2,600 for a azygous tally and progressive AI moving for 19 days without quality intervention.
Difficult, but fair
Reimplementing full programs is highly challenging for quality package engineers. We judge a quality technologist without AI would return months to lick the astir analyzable MirrorCode tasks. However, MirrorCode tasks are besides feasible; we cognize that location is capable accusation for the tasks to beryllium fair.
Cheat-resistant by design
We sandbox AI models, requiring them to behaviour their activity without entree to the internet, without entree to the original codebase, and pinch nary measurement to cheat connected the task. There are end-to-end tests that models ne'er spot while processing their code, truthful they cannot simply create a lookup array to mimic the original program's outputs.
AI tin already execute immoderate long-horizon coding tasks
AI tin already lick long-horizon MirrorCode tasks, contempt their difficulty. For example, Claude Opus 4.7 reimplemented gotree: a bioinformatics toolkit pinch ~16,000 lines of Go and 40+ commands.1 We judge this aforesaid task would return a quality technologist without AI assistance 2–17 weeks. Opus 4.7 solved it successful 14 hours, costing $251.
One important caveat to these results is information contamination. Because MirrorCode tasks impact reimplementing open-source programs, AI models are apt to person seen the original codebases successful pretraining. This mightiness lead to inflated capacity connected the benchmark. However, AI successfully reimplemented respective target programs that passed our mahfuz screen, and grounded to reimplement programs wherever the surface showed grounds of memorization. This suggests that the results were not dominated by memorization, but we cannot norm retired the anticipation that mahfuz contributes to AI performance. Overall, we expect that the capabilities measured by MirrorCode would generalize to an unseen codebase. We talk this further, on pinch much results and specifications connected benchmark construction, successful the paper.
Leaderboard
MirrorCode is not afloat solved. For our regularly updated leaderboard, we study MirrorCode (ML, +Private, 2L). This intends we tally the 15 target programs from the Medium and Large buckets, and driblet the Small bucket. Each target programme is evaluated successful 2 implementation languages (generally Go and Ada) giving 30 tasks. We tally each task 3 times, pinch a fund of 10 cardinal tokens per attempt.2
Open-source code
We merchandise our scaffold and 22 of the 25 MirrorCode target programs (totaling 132 task instances crossed the six supported programming languages) arsenic open-source, pinch the different 3 targets held retired arsenic a backstage trial set.
This activity was co-developed pinch METR and supported by a assistance from METR. The authors of MirrorCode are Tom Adamczewski, David Owen, and David Rein. Florian Brand, Giles Edkins, Allen Hart, and Daniel O’Connell contributed further target programs. Rasmus Faber-Espensen made important infrastructure improvements and gave proposal connected engineering
English (US) ·
Indonesian (ID) ·