Skip to notes
aicharts
Theme
Appearance

Open models closed SemiAnalysis composites, not this table

SemiAnalysis’s faster catch-up describes era composites. Closed configurations still lead AA Index in this coding-agent snapshot.

Drafted with AI and reviewed by Codex.

A dark rounded core sits inside a gold lattice frame with four open loops.
Each result belongs to a complete model, harness, and setting configuration. Made with SlopCamera

SemiAnalysis asks whether open models are catching up, in an essay published August 21, 2026 by Evan Cloutier, Max Kan, Jordan Nanos, and Dylan Patel. Their answer is about era-specific composites: catch-up time has halved with each generation, down to 4.8 to 6 months in the agentic era. This note keeps that claim in its own measurement, then asks a different question of the current coding-agent leaders table.

The Artificial Analysis snapshot stored by aicharts names a model, an agent harness, an effort setting, and a mean API cost for every row. It does not publish an open-versus-closed field. The comparison below uses an explicit provider allowlist, copies scores from the checked snapshot, and quotes only SemiAnalysis figures that appear in the essay text. The two sources can agree that open models have become more useful without agreeing that they have closed the coding-agent table.

SemiAnalysis measures era composites

SemiAnalysis refuses a single historical scoreboard. Early-scaling exams saturate, reasoning exams replace them, and agentic work then needs terminal, browsing, and software-engineering tasks. Their Era 3 suite is Terminal-Bench 2.1, BrowseComp-Plus, τ³-banking, and DeepSWE. They ran most scores on Prime Intellect's evaluation stack and used additional runs from Artificial Analysis and Datacurve.

On that design they report a cycle. A closed lab jumps first. Other labs reverse-engineer the advance, including through distillation, and close the gap. In the early-scaling era their composite is 75.7 for GPT-3.5 Turbo and 39.9 for Llama-2-70B. Llama-3.1-405B later reaches 86. GPT-4o and DeepSeek V3 finish the era at 95.5 and 94.1.

The reasoning-era opening gap is 12.1 points, against 35.8 at the start of the previous era. DeepSeek R1-0528 closes that opening gap at 78 after 8.5 months. In the agentic era they report that Kimi K2.6 surpassed Opus 4.5 at 56.3 in 4.8 months, and that GLM-5.2 cleared GPT-5.2 at 72.4 in 6 months.

Those sentences are SemiAnalysis measurements, not aicharts calculations. The essay also limits what the composites prove. The authors still prefer Fable 5 for daily work over Kimi K3, even while saying Kimi K3 may score higher on their curated suite. They treat public benchmarks as hill-climbable: labs can train reinforcement-learning environments that mimic the evals.

What the coding-agent snapshot records

aicharts retrieved the checked snapshot on Oct 5, 2026, 2:10 PM UTC. The dataset contains 31 model-agent configurations across 22 models, 9 agent harnesses, and 10 providers. AA Index is the snapshot's overall 0–100 score across code changes, terminal work, and repository understanding. DeepSWE v1.1, Terminal-Bench 4, and SWE-Atlas-QnA also have separate columns. The dataset page lists every configuration and the highest stored score for each metric.

This is a closer relative of SemiAnalysis's agentic era than of their earlier exams, but it is not the same composite. The snapshot omits BrowseComp-Plus and τ³-banking, adds SWE-Atlas-QnA, and reports the three component scores alongside their AA Index composite. Every row also carries a harness and setting. A model name without those fields is an incomplete citation here.

Closed configurations still lead on AA Index

The highest AA Index in this snapshot is 68.4 for Sonnet 5.5 on Claude Code at the max setting, with a mean API cost of $14.19 per task. The next stored scores belong to other closed-lab configurations. That is an observation of this table, not a claim that the same systems lead on SemiAnalysis's suite.

Highest AA Index configurations in the Artificial Analysis snapshot retrieved Oct 5, 2026, 2:10 PM UTC
ModelAgentSettingAA IndexCost
Sonnet 5.5Claude Codemax68.4$14.19
Opus 5.5Claude Codemax66.0$13.04
Gemini 4 ArgonAntigravity CLIdefault63.8$5.84
GPT-6.1 SolCodexxhigh62.9$1.04
Sonnet 5.5Claude Codexhigh62.9$3.33
Fable 5.1 (with fallback)Claude Codemax62.2$12.39
Claude Fable 5.1 XHigh + SWE-2 MediumDevin Fusion CLIdefault61.7$7.90
GPT-6 AstraCodexmax61.6$7.47

Open the comparison chart to change axes or pin a model. The leaders table on that page is the same checked snapshot, not a live scrape of the upstream page.

Highest open-weight configurations in this snapshot

aicharts classifies a configuration as open-weight only when its provider is Alibaba Cloud, DeepSeek, Moonshot AI, and Z.ai. Those families are the ones SemiAnalysis treats as open in the essay (DeepSeek, Kimi, Qwen, and GLM). The snapshot does not state a license, so this allowlist is an analysis choice. Cognition and Meta are left unclassified because the snapshot does not state a license and the SemiAnalysis essay does not name those families as open. Cursor, xAI, Google, OpenAI, and Anthropic stay closed.

Under that rule, the highest open-weight AA Index is 53.6 for GLM-5.3 on Opencode at the default setting, with a mean API cost of $4.24 per task. That is 14.8 AA Index points behind Sonnet 5.5 on Claude Code. The gap is aicharts subtraction of two stored scores. It is not a SemiAnalysis composite.

Highest open-weight AA Index configurations in the Artificial Analysis snapshot retrieved Oct 5, 2026, 2:10 PM UTC
ModelAgentProviderSettingAA IndexCost
GLM-5.3OpencodeZ.aidefault53.6$4.24
Kimi K3Kimi Code CLIMoonshot AIdefault51.9$5.05
Qwen3.8 MaxClaude CodeAlibaba Clouddefault43.3$3.48
DeepSeek V4 Pro 0813CodexDeepSeekmax43.1$0.238
DeepSeek V4 Flash 0731CodexDeepSeekmax38.7$0.085

5 open-weight configurations are shown, from 5 classified open-weight rows in the snapshot. Several of those rows use Claude Code or Codex rather than a first-party harness. The snapshot therefore mixes model weights with another lab's agent product. That is one reason a model-only catch-up story and this table can diverge.

The same named models sit in different places

The SemiAnalysis essay quotes agentic-era catch-up scores for named models that are not present in this snapshot. No side-by-side row is possible until those names reappear in a validated refresh.

SemiAnalysis also writes that Kimi K3 may outscore Fable 5 on their composite while they still prefer Fable for daily work. This snapshot does not contain a SemiAnalysis composite for either name, so no Kimi K3-versus-Fable 5 number is quoted from that suite. The available AA Index configurations can be read on the full configuration table.

Cost changes which gap you see

AA Index leaders in this snapshot are expensive relative to the cheapest rows. The AA Index versus cost note keeps a configuration on the frontier when no other row scores at least as high at no greater cost, with a strict improvement on at least one measure. That derived view is aicharts analysis of the stored pairs.

At least one classified open-weight configuration is on that frontier: DeepSeek V4 Flash 0731 on Codex at the max setting, with AA Index 38.7 and mean API cost $0.085 per task. 2 open-weight frontier points appear in this snapshot. Open-weight rows are more visible when the question is inexpensive score than when the question is the highest AA Index.

The useful sentence is narrower than the essay title. Open-weight coding agents in this snapshot are close enough to matter on cost and close enough to appear in the middle of the AA Index list. They are not the current AA Index leaders. SemiAnalysis's faster catch-up time describes their composites, not this table.

Limits of this comparison

  • SemiAnalysis defines and operates its era composites. aicharts does not rerun that suite or recover unpublished chart points from images.
  • Artificial Analysis defines the coding-agent scores and costs. aicharts is an independent visualization and is not affiliated with Artificial Analysis, SemiAnalysis, or the listed providers.
  • The open-weight set is an explicit provider allowlist, not a field in the snapshot. A license change, a new provider, or a different definition of open would change the grouped rows.
  • Scores belong to the named model, harness, setting, task set, and evaluation version on the retrieval date. They do not establish results for every repository or production workflow.
  • This is a checked snapshot, not a live mirror. Cite the retrieval timestamp when quoting a value.

Sources

  1. Are Open Models Catching Up?SemiAnalysis, 2026. The August 21, 2026 essay reports era-specific open-versus-closed composites, catch-up intervals, and the limits of public-benchmark scores.
  2. Coding AgentsArtificial Analysis, 2026. The public coding-agents comparison is the source of the aicharts coding-agent snapshot. Model names, agent harnesses, settings, AA Index scores, and mean API costs are Artificial Analysis measurements.

Figures come from the cited primary sources and the aicharts datasets. aicharts did not rerun the reported benchmarks.