Skip to notes
aicharts
Theme
Appearance

Sonnet 5.5 is first on the coding-agent chart at 68.4, $14.19

Claude Code · Sonnet 5.5 (max) scores 68.4 on AA Index at $14.19 per task, first of 31 configurations and the costliest row. The next frontier point down gives up 2.4 points for 92% of the cost.

Drafted with AI and reviewed by Codex.

Five shallow ivory terraces rise along one low charcoal rail; a matte ivory wedge sits on the highest terrace, with a short brass line on that terrace’s leading edge.
Claude Code · Sonnet 5.5 holds the coding-agent chart’s top score and its highest cost per task at once, and the same harness stores four cheaper settings below it. Made with SlopCamera

In the aicharts coding-agent snapshot retrieved Oct 5, 2026, 2:10 PM UTC, Claude Code · Sonnet 5.5 (max) scores 68.4 on AA Index at a mean API cost of $14.19 per task, first of the 31 configurations that carry an index and the highest cost per task on the chart. The nearest cost-frontier configuration below it, Claude Code · Opus 5.5 (max), gives up 2.4 points for 92% of the cost. The same harness stores five Sonnet 5.5 settings; the headline row is the highest of them. The snapshot’s update log records the row on October 5, 2026.

One model, five settings in one harness

The coding-agent chart is a daily snapshot of the public Artificial Analysis coding-agents comparison. Each row is one configuration: a model, the agent harness that ran it, and an effort setting. Artificial Analysis runs the harness on tasks from three coding benchmarks, DeepSWE v1.1, Terminal-Bench 4, and SWE-Atlas-QnA, and AA Index is the mean of the three scores. The cost is the mean API bill for one task in that harness at list prices, including every tool call and every repeated read of the repository the harness sends.

Claude Sonnet 5.5’s headline row is the model inside Claude Code, Anthropic’s own coding agent, at the max setting. One task in that configuration used 27.7 million tokens and took 87 minutes of harness time on average. The same model at another setting, or in another harness, is another row; the snapshot stores five configurations running Claude Sonnet 5.5: Claude Code · Sonnet 5.5 (max), Claude Code · Sonnet 5.5 (xhigh), Claude Code · Sonnet 5.5 (high), Claude Code · Sonnet 5.5 (medium), and Claude Code · Sonnet 5.5 (low).

First of 31 configurations, at the chart’s highest cost

The 68.4 places Claude Code · Sonnet 5.5 (max) first of the 31 configurations that carry an AA Index in the snapshot retrieved Oct 5, 2026, 2:10 PM UTC. No configuration scores higher. Its $14.19 per task is also the highest cost of the 31 configurations that carry a cost, so the row sits at the top right of the chart as both the highest-scoring configuration and the most expensive to run one task through.

The nearest score below it is Claude Code · Opus 5.5 (max) at 66.0, 2.4 points lower for 92% of the Sonnet 5.5 row’s cost. The table lists the four highest-scoring configurations after it, with each cost as a share of its $14.19.

The four highest-scoring configurations after Claude Code · Sonnet 5.5 (max) in the snapshot retrieved Oct 5, 2026, 2:10 PM UTC
ConfigurationSettingAA IndexCost per taskPoints below Sonnet 5.5Share of Sonnet 5.5’s cost
Claude Code · Opus 5.5max66.0$13.042.4 points92%
Antigravity CLI · Gemini 4 Argondefault63.8$5.844.6 points41%
Codex · GPT-6.1 Solxhigh62.9$1.045.5 points7.3%
Claude Code · Sonnet 5.5xhigh62.9$3.335.5 points23%

What each of the five Claude Code settings buys

The snapshot stores five Claude Code settings for Claude Sonnet 5.5, from low at $0.483 to max at $14.19. The last step, from xhigh to max, adds 5.5 points at 4.3x the cost of the setting below it.

Claude Code · Sonnet 5.5 settings in the snapshot retrieved Oct 5, 2026, 2:10 PM UTC, cheapest first
SettingAA IndexCost per taskPoints over cheaper settingCost multiple over cheaper setting
low42.1$0.483--
medium45.9$0.6193.8 points1.3x
high55.0$1.249.1 points2.0x
xhigh62.9$3.337.9 points2.7x
max68.4$14.195.5 points4.3x

The first cheaper frontier row is Opus 5.5

The chart’s cost frontier contains configurations for which no other row costs no more and scores at least as high, with a strict improvement in either measure. Claude Code · Sonnet 5.5 (max) is on it, at the frontier’s highest score. When configurations tie for that score, a cheaper tied row dominates a more expensive one. The question the frontier answers is what a reader gives up by stepping down from this score to a cheaper row that nothing dominates.

The first step down is Claude Code · Opus 5.5 (max): 2.4 points lower for 92% of the $14.19. The first vertex at half the cost or less is Antigravity CLI · Gemini 4 Argon (default), which gives up 4.6 points for 41% of the cost. The frontier ends at Codex · DeepSeek V4 Flash 0731 (max): 29.6 points lower for 0.6% of the cost.

Cost-frontier configurations below Claude Code · Sonnet 5.5 (max) in the snapshot retrieved Oct 5, 2026, 2:10 PM UTC, highest AA Index first
ConfigurationSettingAA IndexCost per taskPoints below Sonnet 5.5Share of Sonnet 5.5’s cost
Claude Code · Opus 5.5max66.0$13.042.4 points92%
Antigravity CLI · Gemini 4 Argondefault63.8$5.844.6 points41%
Codex · GPT-6.1 Solxhigh62.9$1.045.5 points7.3%
Codex · GPT-6.1 Solmedium61.4$0.7056.9 points5.0%
Codex · GPT-6.1 Sollow57.2$0.49911.1 points3.5%
Codex · GPT-5.6 Lunamax43.2$0.43825.1 points3.1%
Codex · DeepSeek V4 Pro 0813max43.1$0.23825.3 points1.7%
Codex · GPT-6 Lunamax41.1$0.17627.3 points1.2%
Codex · DeepSeek V4 Flash 0731max38.7$0.08529.6 points0.6%

Beside Claude Code · Opus 5.5 at the same setting

The snapshot stores both models in Claude Code at the max setting. Claude Code · Opus 5.5 (max) scores 66.0 at $13.04 per task. Sonnet 5.5 adds 2.4 points at 1.1x the mean cost per task, 1.8x the total tokens per task, and 1.4x the mean time per task. The Claude Code · Opus 5.5 note places that row and compares it with Claude Code · Opus 5.

Claude Code · Sonnet 5.5 and Claude Code · Opus 5.5 at the max setting in the snapshot retrieved Oct 5, 2026, 2:10 PM UTC
MeasureSonnet 5.5Opus 5.5Change
AA Index68.466.0+2.4 points
DeepSWE v1.172.068.4+3.5 points
Terminal-Bench 466.263.1+3.0 points
SWE-Atlas-QnA66.966.4+0.5 points
Mean API cost per task$14.19$13.041.1x
Total tokens per task27.7 million15.6 million1.8x
Mean time per task87 minutes64 minutes1.4x

DeepSWE v1.1 is the component it does not lead

AA Index is the mean of three component benchmarks, and the Sonnet 5.5 row does not lead all three. On DeepSWE v1.1 it scores 72.0, sixth of 31; on Terminal-Bench 4 it scores 66.2, first of 31; and on SWE-Atlas-QnA it scores 66.9, first of 31 configurations that carry each score.

It leads Terminal-Bench 4 by 3.0 points over Claude Code · Opus 5.5 (max) and SWE-Atlas-QnA by 0.5 points over Claude Code · Opus 5.5 (max). Antigravity CLI · Gemini 4 Argon (default) scores 6.8 points higher on DeepSWE v1.1. The composite lead is a Terminal-Bench 4 and SWE-Atlas-QnA lead that the DeepSWE v1.1 gap does not cancel.

Claude Code · Sonnet 5.5 (max) on each AA Index component in the snapshot retrieved Oct 5, 2026, 2:10 PM UTC, against the best other configuration that carries the component
ComponentSonnet 5.5 scoreRankBest other configurationGap
DeepSWE v1.172.06 of 31Antigravity CLI · Gemini 4 Argon (default) at 78.8−6.8 points
Terminal-Bench 466.21 of 31Claude Code · Opus 5.5 (max) at 63.1+3.0 points
SWE-Atlas-QnA66.91 of 31Claude Code · Opus 5.5 (max) at 66.4+0.5 points

Every configuration with a lower AA Index also scores lower on Terminal-Bench 4, so a reader who cares only about terminal work orders the top of this chart the same way the composite does.

The Index score is a different unit

The Artificial Analysis Intelligence Index runs the model through its API under one harness that is the same for every model, across 10 evaluations at version 4.3.2, and its cost per task is the average bill for one of those evaluation tasks. In the snapshot retrieved Oct 6, 2026, 11:15 AM UTC, Claude Sonnet 5.5 (Max, Default Fallback) scores 56.0 at $7.67 per task, second of 103 configurations with a measured cost.

Claude Sonnet 5.5 at max effort in the two aicharts snapshots, each on its own task set with its own cost definition
ChartConfigurationScoreCost per taskSnapshot retrieved
Coding agents (AA Index)Claude Code · Sonnet 5.5 (max)68.4$14.19Oct 5, 2026, 2:10 PM UTC
Intelligence IndexClaude Sonnet 5.5 (Max, Default Fallback)56.0$7.67Oct 6, 2026, 11:15 AM UTC

The 68.4 is a mean of three coding benchmarks run inside Claude Code, and the 56.0 is a weighted average of 10 evaluations run through the API. The $14.19 includes every tool call and every repeated read of the repository that the harness sends, 27.7 million tokens per task in this snapshot; the $7.67 is the average bill for one evaluation task under Artificial Analysis’s standardized harness, with 197,430 output tokens per task.

Limits

  • The chart scores, task costs, token counts, and durations are Artificial Analysis measurements of the named configuration on the retrieval date, under DeepSWE v1.1, Terminal-Bench 4, and SWE-Atlas-QnA for the coding-agent chart and Intelligence Index version 4.3.2 for the capability chart. None of them establishes a result on other repositories, tasks, or harnesses.
  • The rank, cost rank, frontier steps, setting multiples, component gaps, and same-harness multiples are aicharts derivations from the snapshots named in each caption. A configuration added, removed, or rescored by Artificial Analysis moves them, and the coding-agent snapshot advances daily.
  • Claude Sonnet 5.5 in Cursor, Devin, or any harness other than Claude Code is a configuration this snapshot does not store unless a row appears above, so this note says nothing about a missing harness.
  • The Opus 5.5 comparison holds the harness and setting fixed, but the snapshot records outcomes, not run dates or benchmark versions at run time; Artificial Analysis may have measured the two models days apart.
  • The two charts use different task sets and different cost definitions. Their scores are not one ranking.

Sources

  1. Coding AgentsArtificial Analysis, 2026. The public coding-agents comparison is the source of the aicharts coding-agent snapshot. Model names, agent harnesses, settings, AA Index scores, and mean API costs are Artificial Analysis measurements.
  2. LLM LeaderboardArtificial Analysis, 2026. The public models leaderboard is the source of the aicharts Intelligence Index snapshot. Scores, per-task costs, and output tokens are Artificial Analysis measurements under Intelligence Index v4.3.2.

Figures come from the cited primary sources and the aicharts datasets. aicharts did not rerun the reported benchmarks.