Skip to benchmark notes

Are open models catching up on coding-agent benchmarks?

SemiAnalysis reports a shrinking open-versus-closed gap on era-specific composites. The current coding-agent snapshot answers a narrower question: named model, harness, setting, and cost.

SemiAnalysis asks whether open models are catching up, in an essay published August 21, 2026 by Evan Cloutier, Max Kan, Jordan Nanos, and Dylan Patel. Their answer is about era-specific composites: catch-up time has halved with each generation, down to 4.8 to 6 months in the agentic era. This note keeps that claim in its own measurement, then asks a different question of the current coding-agent leaders table.

The Artificial Analysis snapshot stored by AI Charts names a model, an agent harness, an effort setting, and a mean API cost for every row. It does not publish an open-versus-closed field. The comparison below uses an explicit provider allowlist, copies scores from the checked snapshot, and quotes only SemiAnalysis figures that appear in the essay text. The two sources can agree that open models have become more useful without agreeing that they have closed the coding-agent table.

SemiAnalysis measures era composites

SemiAnalysis refuses a single historical scoreboard. Early-scaling exams saturate, reasoning exams replace them, and agentic work then needs terminal, browsing, and software-engineering tasks. Their Era 3 suite is Terminal-Bench 2.1, BrowseComp-Plus, τ³-banking, and DeepSWE. They ran most scores on Prime Intellect's evaluation stack and used additional runs from Artificial Analysis and Datacurve.

On that design they report a cycle. A closed lab jumps first. Other labs reverse-engineer the advance, including through distillation, and close the gap. In the early-scaling era their composite is 75.7 for GPT-3.5 Turbo and 39.9 for Llama-2-70B. Llama-3.1-405B later reaches 86. GPT-4o and DeepSeek V3 finish the era at 95.5 and 94.1.

The reasoning-era opening gap is 12.1 points, against 35.8 at the start of the previous era. DeepSeek R1-0528 closes that opening gap at 78 after 8.5 months. In the agentic era they report that Kimi K2.6 surpassed Opus 4.5 at 56.3 in 4.8 months, and that GLM-5.2 cleared GPT-5.2 at 72.4 in 6 months.

Those sentences are SemiAnalysis measurements, not AI Charts calculations. The essay also limits what the composites prove. The authors still prefer Fable 5 for daily work over Kimi K3, even while saying Kimi K3 may score higher on their curated suite. They treat public benchmarks as hill-climbable: labs can train reinforcement-learning environments that mimic the evals. A Hraness reading note of the SemiAnalysis essay records the same cycle, catch-up intervals, and caveat.

What the coding-agent snapshot records

AI Charts retrieved the checked snapshot on Aug 18, 2026, 11:03 AM UTC. The dataset contains 59 model-agent configurations across 28 models, 9 agent harnesses, and 10 providers. AA Index is the snapshot's overall 0–100 score across code changes, terminal work, and repository understanding. DeepSWE, Terminal-Bench v2, and SWE-Atlas-QnA stay separate. The dataset page lists every configuration and the highest stored score for each metric.

This is a closer relative of SemiAnalysis's agentic era than of their earlier exams, but it is not the same composite. The snapshot uses Terminal-Bench v2 rather than 2.1, omits BrowseComp-Plus and τ³-banking, adds SWE-Atlas-QnA, and reports DeepSWE as its own column instead of folding it into an unpublished average. Every row also carries a harness and setting. A model name without those fields is an incomplete citation here.

Closed configurations still lead on AA Index

The highest AA Index in this snapshot is 66.7 for Opus 5 on Claude Code at the xhigh setting, with a mean API cost of $8.23 per task. The next stored scores belong to other closed-lab configurations. That is an observation of this table, not a claim that the same systems lead on SemiAnalysis's suite.

Highest AA Index configurations in the Artificial Analysis snapshot retrieved Aug 18, 2026, 11:03 AM UTC
ModelAgentSettingAA IndexCost
Opus 5Claude Codexhigh66.7$8.23
GPT-5.6 SolCodexmax66.6$7.08
Fable 5 (with fallback)Claude Codemax65.8$11.70
Opus 5Claude Codemax65.5$8.95
GPT-5.6 SolCodexxhigh65.1$5.24
Grok 4.5Grok Buildhigh64.4$2.59
GPT-5.6 SolCodexhigh64.1$4.14
Opus 5Claude Codehigh63.4$3.80

Open the comparison chart to change axes or pin a model. The leaders table on that page is the same checked snapshot, not a live scrape of the upstream page.

Highest open-weight configurations in this snapshot

AI Charts classifies a configuration as open-weight only when its provider is Alibaba Cloud, DeepSeek, Moonshot AI, and Z.ai. Those families are the ones SemiAnalysis treats as open in the essay (DeepSeek, Kimi, Qwen, and GLM). The snapshot does not state a license, so this allowlist is an analysis choice. Meta is left unclassified because the snapshot does not state a license and the SemiAnalysis essay does not name that family as open. Cursor, xAI, Google, OpenAI, and Anthropic stay closed.

Under that rule, the highest open-weight AA Index is 61.3 for Kimi K3 on Kimi Code CLI at the default setting, with a mean API cost of $3.18 per task. That is 5.4 AA Index points behind Opus 5 on Claude Code. The gap is AI Charts subtraction of two stored scores. It is not a SemiAnalysis composite.

Highest open-weight AA Index configurations in the Artificial Analysis snapshot retrieved Aug 18, 2026, 11:03 AM UTC
ModelAgentProviderSettingAA IndexCost
Kimi K3Kimi Code CLIMoonshot AIdefault61.3$3.18
Qwen3.8 MaxClaude CodeAlibaba Clouddefault56.6$3.86
DeepSeek V4 FlashCodexDeepSeekmax55.5$0.071
GLM-5.2Claude CodeZ.aidefault43.2$6.51
GLM-5.1Claude CodeZ.aidefault36.1$4.33
Qwen3.7 Plus (thinking)Claude CodeAlibaba Clouddefault36.0$6.23
Kimi K2.6Claude CodeMoonshot AIdefault32.6$1.19
DeepSeek V4 ProClaude CodeDeepSeekhigh31.4$0.272

8 open-weight configurations are shown, from 8 classified open-weight rows in the snapshot. Several of those rows use Claude Code or Codex rather than a first-party harness. The snapshot therefore mixes model weights with another lab's agent product. That is one reason a model-only catch-up story and this table can diverge.

The same named models sit in different places

SemiAnalysis's agentic-era catch-up claim names models that also appear in this snapshot. The essay's composite and the stored AA Index are different measurements of those names. The table copies the quoted SemiAnalysis figure beside the highest AA Index row for the same model string.

Named SemiAnalysis catch-up models that also appear in the Artificial Analysis snapshot retrieved Aug 18, 2026, 11:03 AM UTC
ModelSemiAnalysis compositeClosed referenceMonthsAA Index hereAgent here
Kimi K2.656.3Opus 4.54.832.6Claude Code, default
GLM-5.272.4GPT-5.2643.2Claude Code, default

Kimi K2.6 and GLM-5.2 can close SemiAnalysis's Era 3 composite against older closed flags and still sit well below the current AA Index leaders here. That is not a contradiction in one scoreboard. It is two scoreboards. SemiAnalysis compares era-opening closed models to later open releases on their suite. This snapshot compares current named configurations on Artificial Analysis's coding-agent metrics.

SemiAnalysis also writes that Kimi K3 may outscore Fable 5 on their composite while they still prefer Fable for daily work. This snapshot does not contain a SemiAnalysis composite for either name, so no Kimi K3-versus-Fable 5 number is quoted from that suite. On AA Index, the stored Fable 5 and Kimi K3 rows can be read on the full configuration table.

Cost changes which gap you see

AA Index leaders in this snapshot are expensive relative to the cheapest rows. The AA Index versus cost note keeps a configuration on the frontier only when no other configuration is both cheaper and at least as strong. That derived view is AI Charts analysis of the stored pairs.

At least one classified open-weight configuration is on that frontier: DeepSeek V4 Flash on Codex at the max setting, with AA Index 55.5 and mean API cost $0.071 per task. Open-weight rows are more visible when the question is inexpensive score than when the question is the highest AA Index.

How to read both sources

Read the SemiAnalysis essay for the cycle, the catch-up intervals, and the warning that public benchmarks can be imitated. Read the Hraness reading note for a dated digest of those claims. Read the coding-agent leaders table and the checked dataset when you need the current Artificial Analysis configuration list. Read AA Index versus cost for coding agents when the decision is score against mean API cost rather than open versus closed.

The useful sentence is narrower than the essay title. Open-weight coding agents in this snapshot are close enough to matter on cost and close enough to appear in the middle of the AA Index list. They are not the current AA Index leaders. SemiAnalysis's faster catch-up time describes their composites, not this table.

Limits of this comparison

  • SemiAnalysis defines and operates its era composites. AI Charts does not rerun that suite or recover unpublished chart points from images.
  • Artificial Analysis defines the coding-agent scores and costs. AI Charts is an independent visualization and is not affiliated with Artificial Analysis, SemiAnalysis, or the listed providers.
  • The open-weight set is an explicit provider allowlist, not a field in the snapshot. A license change, a new provider, or a different definition of open would change the grouped rows.
  • Scores belong to the named model, harness, setting, task set, and evaluation version on the retrieval date. They do not establish results for every repository or production workflow.
  • This is a checked snapshot, not a live mirror. Cite the retrieval timestamp when quoting a value.

Sources

  1. Are Open Models Catching Up?SemiAnalysis, 2026. The August 21, 2026 essay reports era-specific open-versus-closed composites, catch-up intervals, and the limits of public-benchmark scores.
  2. Hraness reading note: Are Open Models Catching Up?Hraness, 2026. The Hraness reading note is a dated digest of the SemiAnalysis essay, used here as a crawlable companion citation rather than a substitute for the original.
  3. Coding AgentsArtificial Analysis, 2026. The public coding-agents comparison is the upstream source of the checked AI Charts snapshot. Model names, agent harnesses, settings, AA Index scores, and mean API costs are Artificial Analysis measurements.

Results describe the named model, harness, task set, budget, and evaluation version. They do not establish performance on every production repository.