The current aicharts coding-agent comparison is a checked snapshot of the public Artificial Analysis coding-agents page. This note answers one question from that snapshot: which named model, agent harness, and effort settings lead on AA Index, and which of those rows remain undominated once mean API cost per task is included.
aicharts retrieved the snapshot on Oct 5, 2026, 2:10 PM UTC. The dataset contains 31 model-agent configurations across 22 models, 9 agent harnesses, and 10 providers. 31 of those configurations report both an AA Index and a mean API cost. The values below are copied from that snapshot. aicharts does not recalculate Artificial Analysis scores.
What this snapshot measures
AA Index is the snapshot's overall 0–100 score across code changes, terminal work, and repository understanding. Artificial Analysis also reports DeepSWE v1.1, Terminal-Bench 4, and SWE-Atlas-QnA as separate metrics. Those component scores are not combined here. This note uses AA Index because it is the composite the source publishes for the same configuration that also carries a task-level cost.
API cost is the mean API cost in US dollars for the evaluated task configuration. It is not a subscription price, a latency guarantee, or a production invoice. Active time and total token use exist in the same records and are left for the comparison chart.
Each row is a specific combination of model, agent harness, and effort setting. A model name without the harness and setting is an incomplete citation. Two rows that share a model and differ only in setting are different observations.
Highest AA Index configurations
The highest AA Index in this snapshot is 68.4 for Sonnet 5.5 on Claude Code at the max setting, with a mean API cost of $14.19 per task.
| Model | Agent | Setting | AA Index | Cost |
|---|---|---|---|---|
| Sonnet 5.5 | Claude Code | max | 68.4 | $14.19 |
| Opus 5.5 | Claude Code | max | 66.0 | $13.04 |
| Gemini 4 Argon | Antigravity CLI | default | 63.8 | $5.84 |
| GPT-6.1 Sol | Codex | xhigh | 62.9 | $1.04 |
| Sonnet 5.5 | Claude Code | xhigh | 62.9 | $3.33 |
| Fable 5.1 (with fallback) | Claude Code | max | 62.2 | $12.39 |
| Claude Fable 5.1 XHigh + SWE-2 Medium | Devin Fusion CLI | default | 61.7 | $7.90 |
| GPT-6 Astra | Codex | max | 61.6 | $7.47 |
| GPT-6.1 Sol | Codex | medium | 61.4 | $0.705 |
| GPT-6.1 Sol | Codex | high | 60.1 | $0.889 |
These are the highest stored AA Index scores, not a claim that the same systems lead on DeepSWE v1.1, Terminal-Bench 4, or SWE-Atlas-QnA. The dataset page lists the highest available score for each of those metrics separately and includes every configuration in the snapshot.
Cost and AA Index on the frontier
A configuration is on the cost frontier when no other configuration costs no more and scores at least as high, with a strict improvement in either cost or score. Configurations with identical cost and score share a frontier position.
| Model | Agent | Setting | AA Index | Cost |
|---|---|---|---|---|
| DeepSeek V4 Flash 0731 | Codex | max | 38.7 | $0.085 |
| GPT-6 Luna | Codex | max | 41.1 | $0.176 |
| DeepSeek V4 Pro 0813 | Codex | max | 43.1 | $0.238 |
| GPT-5.6 Luna | Codex | max | 43.2 | $0.438 |
| GPT-6.1 Sol | Codex | low | 57.2 | $0.499 |
| GPT-6.1 Sol | Codex | medium | 61.4 | $0.705 |
| GPT-6.1 Sol | Codex | xhigh | 62.9 | $1.04 |
| Gemini 4 Argon | Antigravity CLI | default | 63.8 | $5.84 |
| Opus 5.5 | Claude Code | max | 66.0 | $13.04 |
| Sonnet 5.5 | Claude Code | max | 68.4 | $14.19 |
The frontier in this snapshot has 10 configurations. Moving between distinct frontier points trades a higher mean task cost for a higher AA Index; the size of the score increase varies.
That sequence is aicharts analysis of the stored pairs. Artificial Analysis does not publish a frontier ranking. The frontier can change when the next validated snapshot adds, removes, or reprices a configuration.
AA Index per dollar is a derived view
Dividing AA Index by mean API cost produces a derived ratio. It is not an Artificial Analysis metric. The ratio favors cheap configurations and can rank a low score above a stronger but more expensive run.
| Model | Agent | Setting | AA Index | Cost | AA Index / $ |
|---|---|---|---|---|---|
| DeepSeek V4 Flash 0731 | Codex | max | 38.7 | $0.085 | 455.4 |
| GPT-6 Luna | Codex | max | 41.1 | $0.176 | 233.5 |
| DeepSeek V4 Pro 0813 | Codex | max | 43.1 | $0.238 | 180.9 |
| GPT-6.1 Sol | Codex | low | 57.2 | $0.499 | 114.7 |
| GPT-5.6 Luna | Codex | max | 43.2 | $0.438 | 98.7 |
| GPT-6.1 Sol | Codex | medium | 61.4 | $0.705 | 87.2 |
| Sonnet 5.5 | Claude Code | low | 42.1 | $0.483 | 87.1 |
| Sonnet 5.5 | Claude Code | medium | 45.9 | $0.619 | 74.2 |
| GPT-6.1 Sol | Codex | high | 60.1 | $0.889 | 67.6 |
| GPT-6.1 Sol | Codex | xhigh | 62.9 | $1.04 | 60.5 |
Use the ratio only to find inexpensive configurations that still have a recorded AA Index. A configuration is on the frontier when no other row scores at least as high at no greater cost, with a strict improvement on at least one measure.
When to use this snapshot
Use this note when you need a sourced answer to a cost and quality question on the current coding-agent snapshot. Open the comparison chart to change axes, pin a model, or inspect provider ranges. Open the dataset page for provenance, benchmark definitions, and the full configuration table.
Limits of the comparison
- Artificial Analysis defines and operates the evaluations. aicharts is an independent visualization and is not affiliated with Artificial Analysis or the listed providers.
- Scores and costs belong to the named model, harness, setting, task set, and evaluation version on the retrieval date. They do not establish results for every repository or production workflow.
- Mean task cost is not a price quote. Prompt mix, retry policy, caching, and live API prices can differ from the evaluation.
- AA Index is a composite. A configuration can lead on the index and trail on a component benchmark.
- This is a checked snapshot, not a live mirror. Cite the retrieval timestamp when quoting a value.
