Skip to benchmark notes

Open models can close a scoreboard and still lose the product

SemiAnalysis measures a shrinking open-versus-closed gap on era composites, then still picks a productized closed stack for daily work. A catch-up number is not a reason to collapse named coding-agent rows.

SemiAnalysis asks whether open models are catching up, in an essay published August 21, 2026 by Evan Cloutier, Max Kan, Jordan Nanos, and Dylan Patel. Their measurement is an era-specific composite. Catch-up time has halved with each generation, down to 4.8 to 6 months in the agentic era. In the same essay they still prefer a productized closed stack for daily work, because public composites are hill-climbable. This note keeps that split.

The current coding-agent comparison names a model, an agent harness, and an effort setting on every stored row. A composite catch-up number is not a reason to collapse those rows into “open won.” The named open-weight coding-agent rows stay on their own page. This page asks what SemiAnalysis’s own product preference changes about that headline.

SemiAnalysis measures catch-up by era

The essay refuses one historical scoreboard. Early-scaling exams saturate, reasoning exams replace them, and agentic work then needs terminal, browsing, and software-engineering tasks. Their Era 3 suite is Terminal-Bench 2.1, BrowseComp-Plus, τ³-banking, and DeepSWE. They ran most scores on Prime Intellect’s evaluation stack and used additional runs from Artificial Analysis and Datacurve.

On that design they report a cycle. A closed lab jumps first. Other labs reverse-engineer the advance, including through distillation, and close the gap. Their measured trend is that each generation takes about half as long to catch the first closed model of its era.

Quoted SemiAnalysis catch-up intervals from the August 21, 2026 essay
EraQuoted catch-upClosed reference they name
Reasoning8.5 months to 78The o1-era opening gap, closed by DeepSeek R1-0528
Agentic4.8 months at 56.3Opus 4.5, passed by Kimi K2.6 on their composite
Agentic6 months at 72.4GPT-5.2, cleared by GLM-5.2 on their composite

Those are SemiAnalysis measurements. They report a faster close in Era 3 than in the two earlier eras. The numbers belong to their suite and their choice of era-opening closed flags. They are not AA Index cells, and they are not a verdict on every later closed product.

The same authors still pick a product

The essay’s daily-work clause sits next to the catch-up chart. Kimi K3 scores higher than Fable 5 on their curated composite, but the authors still prefer Fable for daily work. They credit Anthropic’s product work around the model and warn that benchmarks are incomplete proxies for real use.

They already applied that clause inside the closed field. GPT-5.2 scored higher than Opus 4.5 on their suite, yet their account says the user experience did not follow the composite. They treat the complete model-and-harness product as the relevant unit and Anthropic’s stack as the agentic-era default, even when another closed flag leads the same suite.

A Hraness reading note of the SemiAnalysis essay, saved 2026-08-21, records the same limit as a dated digest: “The authors still prefer Anthropic’s productized stack for daily work and treat public benchmarks as hill-climbable, incomplete proxies for real use.” The digest’s product sentence is the one this page keeps in view: “The product, not the leaderboard, still decides daily use.”

Public composites are hill-climbable

SemiAnalysis limits what the composites prove. A lab can build reinforcement-learning environments that resemble published benchmark tasks and optimize against them. A later open release can pass the era-opening closed flags on that suite and still leave the daily product undecided.

Hill-climbing is a property of a published task set, not a property of open weights. A closed lab can train against the same public composite. An open lab can ship weights that look strong on that composite and weak inside another harness, another setting, or a private workflow. The catch-up interval measures how fast the suite moved. It does not measure whether the winning weights are the thing a person should run tomorrow.

A coding-agent row is a product, not a weight class

Hraness defines an agent harness as software that gives a model a place to work: it injects instructions, offers tools, runs an assess-act-reassess loop, and translates across model APIs. SemiAnalysis’s agentic-era clause uses the same split. The composite can move when the weights move. Daily use, in their telling, still follows the productized loop around those weights.

AI Charts retrieved the checked snapshot on Aug 18, 2026, 11:03 AM UTC. The dataset contains 59 model-agent configurations across 28 models, 9 agent harnesses, and 10 providers. The snapshot has no open-versus-closed field. Each row is already a product citation: a model name, a harness name, and an effort setting, with scores copied from Artificial Analysis’s coding-agents comparison.

The highest AA Index in this snapshot is 66.7 for Opus 5 on Claude Code at the xhigh setting. That sentence is complete only when those four fields stay attached. Replacing it with “open won” or “closed won” would drop the harness, the setting, the suite, and the retrieval date.

Current coding-agent leaders in the Artificial Analysis snapshot retrieved Aug 18, 2026, 11:03 AM UTC
BenchmarkModelAgentSettingScore
AA IndexOpus 5Claude Codexhigh66.7
DeepSWEGPT-5.6 SolCodexmax68.7
Terminal-Bench v2GPT-5.6 SolCodexmax87.7
SWE-Atlas-QnAOpus 5Claude Codexhigh54.8

Those leaders are observations of named configurations, not a weight-class rank. Several stored rows use another lab’s harness around a first-party model. The snapshot therefore already mixes weights with someone else’s agent product. That is one reason a model-only catch-up story and this table can diverge without either source being wrong.

A catch-up number does not collapse the snapshot

SemiAnalysis’s 4.8-to-6-month result compares later open releases with the closed flags that opened the agentic era on their suite. The live snapshot compares current named configurations on Artificial Analysis’s coding-agent metrics. Passing Opus 4.5 or GPT-5.2 on an unpublished average is a different event from leading the checked dataset today.

Open models on coding-agent benchmarks is the page that copies those named harness rows and places SemiAnalysis’s quoted catch-up figures beside matching model strings. That page answers whether classified open-weight rows sit with the current AA Index leaders. This page answers whether the essay’s catch-up number is a reason to stop naming the rows.

What each source is allowed to decide
SourceQuestion it answersWhat a win there means
SemiAnalysis Era 3 compositeDid a later open release pass the era-opening closed flags on their suite?Catch-up on that composite, in the quoted months
AI Charts coding-agent snapshotWhat did this named model, harness, and setting score on the stored metrics?A configuration observation on the retrieval date
SemiAnalysis daily-work noteWhich stack do the authors still use?A product preference, not a second composite

The useful failure mode is a headline that treats the first row as a substitute for the second and third. “Open caught up in 4.8 months” can be a faithful citation of SemiAnalysis and still be the wrong instruction for this snapshot. The snapshot would have to drop harness and setting to print that headline. SemiAnalysis’s own Fable preference is evidence that they do not make that drop either.

How to read the two pages

Read the SemiAnalysis essay for the cycle, the catch-up intervals, the hill-climb limit, and the daily-work preference. Read the Hraness reading note for a dated digest of those claims. Read open models on coding-agent benchmarks when you need the current classified open-weight AA Index cells. Read the coding-agent leaders table and the checked dataset when you need the full configuration list. Read the Hraness harness definition when a row’s agent name needs a noun.

Read why a high score still needs a holdout when the next question is unseen tasks rather than an unseen product layer. Read why coding-agent scores still need expertise when the next question is whether the person citing the cell can specify and audit the work. Those notes share the leaders table. They do not share this product-versus-scoreboard question.

The useful sentence is narrower than the essay title. Open models can close an era composite against the closed flags that opened that era. SemiAnalysis still picks a productized closed stack for daily work. AI Charts keeps the live coding-agent snapshot as named model, harness, and setting rows for the same reason.

Limits of this comparison

  • SemiAnalysis defines and operates its era composites and states its own daily-work preference. AI Charts does not rerun that suite or recover unpublished chart points from images.
  • The Hraness page is a dated digest, not a substitute for the essay. Quote SemiAnalysis for the measurements and the Hraness note only for its own digest sentences.
  • Artificial Analysis defines the coding-agent scores. AI Charts is an independent visualization and is not affiliated with Artificial Analysis, SemiAnalysis, or the listed providers.
  • This page does not classify snapshot rows as open or closed. Weight-class grouping stays on the named-row note, where the allowlist is an explicit analysis choice.
  • Scores belong to the named model, harness, setting, task set, and evaluation version on the retrieval date. They do not establish results for every repository or production workflow.
  • This is a checked snapshot, not a live mirror. Cite the retrieval timestamp when quoting a value.

Sources

  1. Are Open Models Catching Up?SemiAnalysis, 2026. The August 21, 2026 essay reports era-specific open-versus-closed composites, catch-up intervals, and the limits of public-benchmark scores.
  2. Hraness reading note: Are Open Models Catching Up?Hraness, 2026. The Hraness reading note is a dated digest of the SemiAnalysis essay, used here as a crawlable companion citation rather than a substitute for the original.
  3. Coding AgentsArtificial Analysis, 2026. The public coding-agents comparison is the upstream source of the checked AI Charts snapshot. Model names, agent harnesses, settings, AA Index scores, and mean API costs are Artificial Analysis measurements.

Results describe the named model, harness, task set, budget, and evaluation version. They do not establish performance on every production repository.