In the aicharts coding-agent snapshot retrieved Oct 5, 2026, 2:10 PM UTC, Claude Code · Sonnet 5.5 (max) scores 68.4 on AA Index at a mean API cost of $14.19 per task, first of the 31 configurations that carry an index and the highest cost per task on the chart. The nearest cost-frontier configuration below it, Claude Code · Opus 5.5 (max), gives up 2.4 points for 92% of the cost. The same harness stores five Sonnet 5.5 settings; the headline row is the highest of them. The snapshot’s update log records the row on October 5, 2026.
One model, five settings in one harness
The coding-agent chart is a daily snapshot of the public Artificial Analysis coding-agents comparison. Each row is one configuration: a model, the agent harness that ran it, and an effort setting. Artificial Analysis runs the harness on tasks from three coding benchmarks, DeepSWE v1.1, Terminal-Bench 4, and SWE-Atlas-QnA, and AA Index is the mean of the three scores. The cost is the mean API bill for one task in that harness at list prices, including every tool call and every repeated read of the repository the harness sends.
Claude Sonnet 5.5’s headline row is the model inside Claude Code, Anthropic’s own coding agent, at the max setting. One task in that configuration used 27.7 million tokens and took 87 minutes of harness time on average. The same model at another setting, or in another harness, is another row; the snapshot stores five configurations running Claude Sonnet 5.5: Claude Code · Sonnet 5.5 (max), Claude Code · Sonnet 5.5 (xhigh), Claude Code · Sonnet 5.5 (high), Claude Code · Sonnet 5.5 (medium), and Claude Code · Sonnet 5.5 (low).
First of 31 configurations, at the chart’s highest cost
The 68.4 places Claude Code · Sonnet 5.5 (max) first of the 31 configurations that carry an AA Index in the snapshot retrieved Oct 5, 2026, 2:10 PM UTC. No configuration scores higher. Its $14.19 per task is also the highest cost of the 31 configurations that carry a cost, so the row sits at the top right of the chart as both the highest-scoring configuration and the most expensive to run one task through.
The nearest score below it is Claude Code · Opus 5.5 (max) at 66.0, 2.4 points lower for 92% of the Sonnet 5.5 row’s cost. The table lists the four highest-scoring configurations after it, with each cost as a share of its $14.19.
| Configuration | Setting | AA Index | Cost per task | Points below Sonnet 5.5 | Share of Sonnet 5.5’s cost |
|---|---|---|---|---|---|
| Claude Code · Opus 5.5 | max | 66.0 | $13.04 | 2.4 points | 92% |
| Antigravity CLI · Gemini 4 Argon | default | 63.8 | $5.84 | 4.6 points | 41% |
| Codex · GPT-6.1 Sol | xhigh | 62.9 | $1.04 | 5.5 points | 7.3% |
| Claude Code · Sonnet 5.5 | xhigh | 62.9 | $3.33 | 5.5 points | 23% |
What each of the five Claude Code settings buys
The snapshot stores five Claude Code settings for Claude Sonnet 5.5, from low at $0.483 to max at $14.19. The last step, from xhigh to max, adds 5.5 points at 4.3x the cost of the setting below it.
| Setting | AA Index | Cost per task | Points over cheaper setting | Cost multiple over cheaper setting |
|---|---|---|---|---|
| low | 42.1 | $0.483 | - | - |
| medium | 45.9 | $0.619 | 3.8 points | 1.3x |
| high | 55.0 | $1.24 | 9.1 points | 2.0x |
| xhigh | 62.9 | $3.33 | 7.9 points | 2.7x |
| max | 68.4 | $14.19 | 5.5 points | 4.3x |
The first cheaper frontier row is Opus 5.5
The chart’s cost frontier contains configurations for which no other row costs no more and scores at least as high, with a strict improvement in either measure. Claude Code · Sonnet 5.5 (max) is on it, at the frontier’s highest score. When configurations tie for that score, a cheaper tied row dominates a more expensive one. The question the frontier answers is what a reader gives up by stepping down from this score to a cheaper row that nothing dominates.
The first step down is Claude Code · Opus 5.5 (max): 2.4 points lower for 92% of the $14.19. The first vertex at half the cost or less is Antigravity CLI · Gemini 4 Argon (default), which gives up 4.6 points for 41% of the cost. The frontier ends at Codex · DeepSeek V4 Flash 0731 (max): 29.6 points lower for 0.6% of the cost.
| Configuration | Setting | AA Index | Cost per task | Points below Sonnet 5.5 | Share of Sonnet 5.5’s cost |
|---|---|---|---|---|---|
| Claude Code · Opus 5.5 | max | 66.0 | $13.04 | 2.4 points | 92% |
| Antigravity CLI · Gemini 4 Argon | default | 63.8 | $5.84 | 4.6 points | 41% |
| Codex · GPT-6.1 Sol | xhigh | 62.9 | $1.04 | 5.5 points | 7.3% |
| Codex · GPT-6.1 Sol | medium | 61.4 | $0.705 | 6.9 points | 5.0% |
| Codex · GPT-6.1 Sol | low | 57.2 | $0.499 | 11.1 points | 3.5% |
| Codex · GPT-5.6 Luna | max | 43.2 | $0.438 | 25.1 points | 3.1% |
| Codex · DeepSeek V4 Pro 0813 | max | 43.1 | $0.238 | 25.3 points | 1.7% |
| Codex · GPT-6 Luna | max | 41.1 | $0.176 | 27.3 points | 1.2% |
| Codex · DeepSeek V4 Flash 0731 | max | 38.7 | $0.085 | 29.6 points | 0.6% |
Beside Claude Code · Opus 5.5 at the same setting
The snapshot stores both models in Claude Code at the max setting. Claude Code · Opus 5.5 (max) scores 66.0 at $13.04 per task. Sonnet 5.5 adds 2.4 points at 1.1x the mean cost per task, 1.8x the total tokens per task, and 1.4x the mean time per task. The Claude Code · Opus 5.5 note places that row and compares it with Claude Code · Opus 5.
| Measure | Sonnet 5.5 | Opus 5.5 | Change |
|---|---|---|---|
| AA Index | 68.4 | 66.0 | +2.4 points |
| DeepSWE v1.1 | 72.0 | 68.4 | +3.5 points |
| Terminal-Bench 4 | 66.2 | 63.1 | +3.0 points |
| SWE-Atlas-QnA | 66.9 | 66.4 | +0.5 points |
| Mean API cost per task | $14.19 | $13.04 | 1.1x |
| Total tokens per task | 27.7 million | 15.6 million | 1.8x |
| Mean time per task | 87 minutes | 64 minutes | 1.4x |
DeepSWE v1.1 is the component it does not lead
AA Index is the mean of three component benchmarks, and the Sonnet 5.5 row does not lead all three. On DeepSWE v1.1 it scores 72.0, sixth of 31; on Terminal-Bench 4 it scores 66.2, first of 31; and on SWE-Atlas-QnA it scores 66.9, first of 31 configurations that carry each score.
It leads Terminal-Bench 4 by 3.0 points over Claude Code · Opus 5.5 (max) and SWE-Atlas-QnA by 0.5 points over Claude Code · Opus 5.5 (max). Antigravity CLI · Gemini 4 Argon (default) scores 6.8 points higher on DeepSWE v1.1. The composite lead is a Terminal-Bench 4 and SWE-Atlas-QnA lead that the DeepSWE v1.1 gap does not cancel.
| Component | Sonnet 5.5 score | Rank | Best other configuration | Gap |
|---|---|---|---|---|
| DeepSWE v1.1 | 72.0 | 6 of 31 | Antigravity CLI · Gemini 4 Argon (default) at 78.8 | −6.8 points |
| Terminal-Bench 4 | 66.2 | 1 of 31 | Claude Code · Opus 5.5 (max) at 63.1 | +3.0 points |
| SWE-Atlas-QnA | 66.9 | 1 of 31 | Claude Code · Opus 5.5 (max) at 66.4 | +0.5 points |
Every configuration with a lower AA Index also scores lower on Terminal-Bench 4, so a reader who cares only about terminal work orders the top of this chart the same way the composite does.
The Index score is a different unit
The Artificial Analysis Intelligence Index runs the model through its API under one harness that is the same for every model, across 10 evaluations at version 4.3.2, and its cost per task is the average bill for one of those evaluation tasks. In the snapshot retrieved Oct 6, 2026, 11:15 AM UTC, Claude Sonnet 5.5 (Max, Default Fallback) scores 56.0 at $7.67 per task, second of 103 configurations with a measured cost.
| Chart | Configuration | Score | Cost per task | Snapshot retrieved |
|---|---|---|---|---|
| Coding agents (AA Index) | Claude Code · Sonnet 5.5 (max) | 68.4 | $14.19 | Oct 5, 2026, 2:10 PM UTC |
| Intelligence Index | Claude Sonnet 5.5 (Max, Default Fallback) | 56.0 | $7.67 | Oct 6, 2026, 11:15 AM UTC |
The 68.4 is a mean of three coding benchmarks run inside Claude Code, and the 56.0 is a weighted average of 10 evaluations run through the API. The $14.19 includes every tool call and every repeated read of the repository that the harness sends, 27.7 million tokens per task in this snapshot; the $7.67 is the average bill for one evaluation task under Artificial Analysis’s standardized harness, with 197,430 output tokens per task.
Limits
- The chart scores, task costs, token counts, and durations are Artificial Analysis measurements of the named configuration on the retrieval date, under DeepSWE v1.1, Terminal-Bench 4, and SWE-Atlas-QnA for the coding-agent chart and Intelligence Index version 4.3.2 for the capability chart. None of them establishes a result on other repositories, tasks, or harnesses.
- The rank, cost rank, frontier steps, setting multiples, component gaps, and same-harness multiples are aicharts derivations from the snapshots named in each caption. A configuration added, removed, or rescored by Artificial Analysis moves them, and the coding-agent snapshot advances daily.
- Claude Sonnet 5.5 in Cursor, Devin, or any harness other than Claude Code is a configuration this snapshot does not store unless a row appears above, so this note says nothing about a missing harness.
- The Opus 5.5 comparison holds the harness and setting fixed, but the snapshot records outcomes, not run dates or benchmark versions at run time; Artificial Analysis may have measured the two models days apart.
- The two charts use different task sets and different cost definitions. Their scores are not one ranking.
