SemiAnalysis asks whether open models are catching up, in an essay published August 21, 2026 by Evan Cloutier, Max Kan, Jordan Nanos, and Dylan Patel. Their answer is about era-specific composites: catch-up time has halved with each generation, down to 4.8 to 6 months in the agentic era. This note keeps that claim in its own measurement, then asks a different question of the current coding-agent leaders table.
The Artificial Analysis snapshot stored by AI Charts names a model, an agent harness, an effort setting, and a mean API cost for every row. It does not publish an open-versus-closed field. The comparison below uses an explicit provider allowlist, copies scores from the checked snapshot, and quotes only SemiAnalysis figures that appear in the essay text. The two sources can agree that open models have become more useful without agreeing that they have closed the coding-agent table.
SemiAnalysis measures era composites
SemiAnalysis refuses a single historical scoreboard. Early-scaling exams saturate, reasoning exams replace them, and agentic work then needs terminal, browsing, and software-engineering tasks. Their Era 3 suite is Terminal-Bench 2.1, BrowseComp-Plus, τ³-banking, and DeepSWE. They ran most scores on Prime Intellect's evaluation stack and used additional runs from Artificial Analysis and Datacurve.
On that design they report a cycle. A closed lab jumps first. Other labs reverse-engineer the advance, including through distillation, and close the gap. In the early-scaling era their composite is 75.7 for GPT-3.5 Turbo and 39.9 for Llama-2-70B. Llama-3.1-405B later reaches 86. GPT-4o and DeepSeek V3 finish the era at 95.5 and 94.1.
The reasoning-era opening gap is 12.1 points, against 35.8 at the start of the previous era. DeepSeek R1-0528 closes that opening gap at 78 after 8.5 months. In the agentic era they report that Kimi K2.6 surpassed Opus 4.5 at 56.3 in 4.8 months, and that GLM-5.2 cleared GPT-5.2 at 72.4 in 6 months.
Those sentences are SemiAnalysis measurements, not AI Charts calculations. The essay also limits what the composites prove. The authors still prefer Fable 5 for daily work over Kimi K3, even while saying Kimi K3 may score higher on their curated suite. They treat public benchmarks as hill-climbable: labs can train reinforcement-learning environments that mimic the evals. A Hraness reading note of the SemiAnalysis essay records the same cycle, catch-up intervals, and caveat.
What the coding-agent snapshot records
AI Charts retrieved the checked snapshot on Aug 18, 2026, 11:03 AM UTC. The dataset contains 59 model-agent configurations across 28 models, 9 agent harnesses, and 10 providers. AA Index is the snapshot's overall 0–100 score across code changes, terminal work, and repository understanding. DeepSWE, Terminal-Bench v2, and SWE-Atlas-QnA stay separate. The dataset page lists every configuration and the highest stored score for each metric.
This is a closer relative of SemiAnalysis's agentic era than of their earlier exams, but it is not the same composite. The snapshot uses Terminal-Bench v2 rather than 2.1, omits BrowseComp-Plus and τ³-banking, adds SWE-Atlas-QnA, and reports DeepSWE as its own column instead of folding it into an unpublished average. Every row also carries a harness and setting. A model name without those fields is an incomplete citation here.
Closed configurations still lead on AA Index
The highest AA Index in this snapshot is 66.7 for Opus 5 on Claude Code at the xhigh setting, with a mean API cost of $8.23 per task. The next stored scores belong to other closed-lab configurations. That is an observation of this table, not a claim that the same systems lead on SemiAnalysis's suite.
| Model | Agent | Setting | AA Index | Cost |
|---|---|---|---|---|
| Opus 5 | Claude Code | xhigh | 66.7 | $8.23 |
| GPT-5.6 Sol | Codex | max | 66.6 | $7.08 |
| Fable 5 (with fallback) | Claude Code | max | 65.8 | $11.70 |
| Opus 5 | Claude Code | max | 65.5 | $8.95 |
| GPT-5.6 Sol | Codex | xhigh | 65.1 | $5.24 |
| Grok 4.5 | Grok Build | high | 64.4 | $2.59 |
| GPT-5.6 Sol | Codex | high | 64.1 | $4.14 |
| Opus 5 | Claude Code | high | 63.4 | $3.80 |
Open the comparison chart to change axes or pin a model. The leaders table on that page is the same checked snapshot, not a live scrape of the upstream page.
Highest open-weight configurations in this snapshot
AI Charts classifies a configuration as open-weight only when its provider is Alibaba Cloud, DeepSeek, Moonshot AI, and Z.ai. Those families are the ones SemiAnalysis treats as open in the essay (DeepSeek, Kimi, Qwen, and GLM). The snapshot does not state a license, so this allowlist is an analysis choice. Meta is left unclassified because the snapshot does not state a license and the SemiAnalysis essay does not name that family as open. Cursor, xAI, Google, OpenAI, and Anthropic stay closed.
Under that rule, the highest open-weight AA Index is 61.3 for Kimi K3 on Kimi Code CLI at the default setting, with a mean API cost of $3.18 per task. That is 5.4 AA Index points behind Opus 5 on Claude Code. The gap is AI Charts subtraction of two stored scores. It is not a SemiAnalysis composite.
| Model | Agent | Provider | Setting | AA Index | Cost |
|---|---|---|---|---|---|
| Kimi K3 | Kimi Code CLI | Moonshot AI | default | 61.3 | $3.18 |
| Qwen3.8 Max | Claude Code | Alibaba Cloud | default | 56.6 | $3.86 |
| DeepSeek V4 Flash | Codex | DeepSeek | max | 55.5 | $0.071 |
| GLM-5.2 | Claude Code | Z.ai | default | 43.2 | $6.51 |
| GLM-5.1 | Claude Code | Z.ai | default | 36.1 | $4.33 |
| Qwen3.7 Plus (thinking) | Claude Code | Alibaba Cloud | default | 36.0 | $6.23 |
| Kimi K2.6 | Claude Code | Moonshot AI | default | 32.6 | $1.19 |
| DeepSeek V4 Pro | Claude Code | DeepSeek | high | 31.4 | $0.272 |
8 open-weight configurations are shown, from 8 classified open-weight rows in the snapshot. Several of those rows use Claude Code or Codex rather than a first-party harness. The snapshot therefore mixes model weights with another lab's agent product. That is one reason a model-only catch-up story and this table can diverge.
The same named models sit in different places
SemiAnalysis's agentic-era catch-up claim names models that also appear in this snapshot. The essay's composite and the stored AA Index are different measurements of those names. The table copies the quoted SemiAnalysis figure beside the highest AA Index row for the same model string.
| Model | SemiAnalysis composite | Closed reference | Months | AA Index here | Agent here |
|---|---|---|---|---|---|
| Kimi K2.6 | 56.3 | Opus 4.5 | 4.8 | 32.6 | Claude Code, default |
| GLM-5.2 | 72.4 | GPT-5.2 | 6 | 43.2 | Claude Code, default |
Kimi K2.6 and GLM-5.2 can close SemiAnalysis's Era 3 composite against older closed flags and still sit well below the current AA Index leaders here. That is not a contradiction in one scoreboard. It is two scoreboards. SemiAnalysis compares era-opening closed models to later open releases on their suite. This snapshot compares current named configurations on Artificial Analysis's coding-agent metrics.
SemiAnalysis also writes that Kimi K3 may outscore Fable 5 on their composite while they still prefer Fable for daily work. This snapshot does not contain a SemiAnalysis composite for either name, so no Kimi K3-versus-Fable 5 number is quoted from that suite. On AA Index, the stored Fable 5 and Kimi K3 rows can be read on the full configuration table.
Cost changes which gap you see
AA Index leaders in this snapshot are expensive relative to the cheapest rows. The AA Index versus cost note keeps a configuration on the frontier only when no other configuration is both cheaper and at least as strong. That derived view is AI Charts analysis of the stored pairs.
At least one classified open-weight configuration is on that frontier: DeepSeek V4 Flash on Codex at the max setting, with AA Index 55.5 and mean API cost $0.071 per task. Open-weight rows are more visible when the question is inexpensive score than when the question is the highest AA Index.
How to read both sources
Read the SemiAnalysis essay for the cycle, the catch-up intervals, and the warning that public benchmarks can be imitated. Read the Hraness reading note for a dated digest of those claims. Read the coding-agent leaders table and the checked dataset when you need the current Artificial Analysis configuration list. Read AA Index versus cost for coding agents when the decision is score against mean API cost rather than open versus closed.
The useful sentence is narrower than the essay title. Open-weight coding agents in this snapshot are close enough to matter on cost and close enough to appear in the middle of the AA Index list. They are not the current AA Index leaders. SemiAnalysis's faster catch-up time describes their composites, not this table.
Limits of this comparison
- SemiAnalysis defines and operates its era composites. AI Charts does not rerun that suite or recover unpublished chart points from images.
- Artificial Analysis defines the coding-agent scores and costs. AI Charts is an independent visualization and is not affiliated with Artificial Analysis, SemiAnalysis, or the listed providers.
- The open-weight set is an explicit provider allowlist, not a field in the snapshot. A license change, a new provider, or a different definition of open would change the grouped rows.
- Scores belong to the named model, harness, setting, task set, and evaluation version on the retrieval date. They do not establish results for every repository or production workflow.
- This is a checked snapshot, not a live mirror. Cite the retrieval timestamp when quoting a value.