SemiAnalysis asks whether open models are catching up, in an essay published August 21, 2026 by Evan Cloutier, Max Kan, Jordan Nanos, and Dylan Patel. Their measurement is an era-specific composite. Catch-up time has halved with each generation, down to 4.8 to 6 months in the agentic era. In the same essay they still prefer a productized closed stack for daily work, because public composites are hill-climbable. This note keeps that split.
The current coding-agent comparison names a model, an agent harness, and an effort setting on every stored row. A composite catch-up number is not a reason to collapse those rows into “open won.” The named open-weight coding-agent rows stay on their own page. This page asks what SemiAnalysis’s own product preference changes about that headline.
SemiAnalysis measures catch-up by era
The essay refuses one historical scoreboard. Early-scaling exams saturate, reasoning exams replace them, and agentic work then needs terminal, browsing, and software-engineering tasks. Their Era 3 suite is Terminal-Bench 2.1, BrowseComp-Plus, τ³-banking, and DeepSWE. They ran most scores on Prime Intellect’s evaluation stack and used additional runs from Artificial Analysis and Datacurve.
On that design they report a cycle. A closed lab jumps first. Other labs reverse-engineer the advance, including through distillation, and close the gap. Their measured trend is that each generation takes about half as long to catch the first closed model of its era.
| Era | Quoted catch-up | Closed reference they name |
|---|---|---|
| Reasoning | 8.5 months to 78 | The o1-era opening gap, closed by DeepSeek R1-0528 |
| Agentic | 4.8 months at 56.3 | Opus 4.5, passed by Kimi K2.6 on their composite |
| Agentic | 6 months at 72.4 | GPT-5.2, cleared by GLM-5.2 on their composite |
Those are SemiAnalysis measurements. They report a faster close in Era 3 than in the two earlier eras. The numbers belong to their suite and their choice of era-opening closed flags. They are not AA Index cells, and they are not a verdict on every later closed product.
The same authors still pick a product
The essay’s daily-work clause sits next to the catch-up chart. Kimi K3 scores higher than Fable 5 on their curated composite, but the authors still prefer Fable for daily work. They credit Anthropic’s product work around the model and warn that benchmarks are incomplete proxies for real use.
They already applied that clause inside the closed field. GPT-5.2 scored higher than Opus 4.5 on their suite, yet their account says the user experience did not follow the composite. They treat the complete model-and-harness product as the relevant unit and Anthropic’s stack as the agentic-era default, even when another closed flag leads the same suite.
A Hraness reading note of the SemiAnalysis essay, saved 2026-08-21, records the same limit as a dated digest: “The authors still prefer Anthropic’s productized stack for daily work and treat public benchmarks as hill-climbable, incomplete proxies for real use.” The digest’s product sentence is the one this page keeps in view: “The product, not the leaderboard, still decides daily use.”
Public composites are hill-climbable
SemiAnalysis limits what the composites prove. A lab can build reinforcement-learning environments that resemble published benchmark tasks and optimize against them. A later open release can pass the era-opening closed flags on that suite and still leave the daily product undecided.
Hill-climbing is a property of a published task set, not a property of open weights. A closed lab can train against the same public composite. An open lab can ship weights that look strong on that composite and weak inside another harness, another setting, or a private workflow. The catch-up interval measures how fast the suite moved. It does not measure whether the winning weights are the thing a person should run tomorrow.
A coding-agent row is a product, not a weight class
Hraness defines an agent harness as software that gives a model a place to work: it injects instructions, offers tools, runs an assess-act-reassess loop, and translates across model APIs. SemiAnalysis’s agentic-era clause uses the same split. The composite can move when the weights move. Daily use, in their telling, still follows the productized loop around those weights.
AI Charts retrieved the checked snapshot on Aug 18, 2026, 11:03 AM UTC. The dataset contains 59 model-agent configurations across 28 models, 9 agent harnesses, and 10 providers. The snapshot has no open-versus-closed field. Each row is already a product citation: a model name, a harness name, and an effort setting, with scores copied from Artificial Analysis’s coding-agents comparison.
The highest AA Index in this snapshot is 66.7 for Opus 5 on Claude Code at the xhigh setting. That sentence is complete only when those four fields stay attached. Replacing it with “open won” or “closed won” would drop the harness, the setting, the suite, and the retrieval date.
| Benchmark | Model | Agent | Setting | Score |
|---|---|---|---|---|
| AA Index | Opus 5 | Claude Code | xhigh | 66.7 |
| DeepSWE | GPT-5.6 Sol | Codex | max | 68.7 |
| Terminal-Bench v2 | GPT-5.6 Sol | Codex | max | 87.7 |
| SWE-Atlas-QnA | Opus 5 | Claude Code | xhigh | 54.8 |
Those leaders are observations of named configurations, not a weight-class rank. Several stored rows use another lab’s harness around a first-party model. The snapshot therefore already mixes weights with someone else’s agent product. That is one reason a model-only catch-up story and this table can diverge without either source being wrong.
A catch-up number does not collapse the snapshot
SemiAnalysis’s 4.8-to-6-month result compares later open releases with the closed flags that opened the agentic era on their suite. The live snapshot compares current named configurations on Artificial Analysis’s coding-agent metrics. Passing Opus 4.5 or GPT-5.2 on an unpublished average is a different event from leading the checked dataset today.
Open models on coding-agent benchmarks is the page that copies those named harness rows and places SemiAnalysis’s quoted catch-up figures beside matching model strings. That page answers whether classified open-weight rows sit with the current AA Index leaders. This page answers whether the essay’s catch-up number is a reason to stop naming the rows.
| Source | Question it answers | What a win there means |
|---|---|---|
| SemiAnalysis Era 3 composite | Did a later open release pass the era-opening closed flags on their suite? | Catch-up on that composite, in the quoted months |
| AI Charts coding-agent snapshot | What did this named model, harness, and setting score on the stored metrics? | A configuration observation on the retrieval date |
| SemiAnalysis daily-work note | Which stack do the authors still use? | A product preference, not a second composite |
The useful failure mode is a headline that treats the first row as a substitute for the second and third. “Open caught up in 4.8 months” can be a faithful citation of SemiAnalysis and still be the wrong instruction for this snapshot. The snapshot would have to drop harness and setting to print that headline. SemiAnalysis’s own Fable preference is evidence that they do not make that drop either.
How to read the two pages
Read the SemiAnalysis essay for the cycle, the catch-up intervals, the hill-climb limit, and the daily-work preference. Read the Hraness reading note for a dated digest of those claims. Read open models on coding-agent benchmarks when you need the current classified open-weight AA Index cells. Read the coding-agent leaders table and the checked dataset when you need the full configuration list. Read the Hraness harness definition when a row’s agent name needs a noun.
Read why a high score still needs a holdout when the next question is unseen tasks rather than an unseen product layer. Read why coding-agent scores still need expertise when the next question is whether the person citing the cell can specify and audit the work. Those notes share the leaders table. They do not share this product-versus-scoreboard question.
The useful sentence is narrower than the essay title. Open models can close an era composite against the closed flags that opened that era. SemiAnalysis still picks a productized closed stack for daily work. AI Charts keeps the live coding-agent snapshot as named model, harness, and setting rows for the same reason.
Limits of this comparison
- SemiAnalysis defines and operates its era composites and states its own daily-work preference. AI Charts does not rerun that suite or recover unpublished chart points from images.
- The Hraness page is a dated digest, not a substitute for the essay. Quote SemiAnalysis for the measurements and the Hraness note only for its own digest sentences.
- Artificial Analysis defines the coding-agent scores. AI Charts is an independent visualization and is not affiliated with Artificial Analysis, SemiAnalysis, or the listed providers.
- This page does not classify snapshot rows as open or closed. Weight-class grouping stays on the named-row note, where the allowlist is an explicit analysis choice.
- Scores belong to the named model, harness, setting, task set, and evaluation version on the retrieval date. They do not establish results for every repository or production workflow.
- This is a checked snapshot, not a live mirror. Cite the retrieval timestamp when quoting a value.