MirrorCode: how far can coding agents work on their own?
MirrorCode tests whether a coding agent can reimplement a complete program under strict end-to-end tests and project-scale resource budgets.
Sourced analysis of AI model and agent benchmarks: what each evaluation measures, what the results show, and where the evidence stops. The first collection focuses on coding agents.
Explore the coding-agent chart2 sourced benchmark summaries
MirrorCode tests whether a coding agent can reimplement a complete program under strict end-to-end tests and project-scale resource budgets.
SlopCodeBench follows agents as they repeatedly extend their own code, measuring correctness, cost, structural erosion, and verbosity at each checkpoint.
Each note starts with the benchmark paper or maintained source page. Material claims link to those primary sources.
Leaderboard values are paired with their observation date and named configuration. They can change after publication.
Methodology limits and interpretation are kept near the results they qualify. Benchmark performance is not treated as a general production claim.