Skip to benchmark notes

AA Index versus cost for coding agents

The checked Artificial Analysis snapshot shows which coding-agent configurations lead on AA Index and how those scores trade off against mean API cost per task.

The current AI Charts coding-agent comparison is a checked snapshot of the public Artificial Analysis coding-agents page. This note answers one question from that snapshot: which named model, agent harness, and effort settings lead on AA Index, and what mean API cost per task those configurations report.

AI Charts retrieved the snapshot on Aug 18, 2026, 11:03 AM UTC. The dataset contains 59 model-agent configurations across 28 models, 9 agent harnesses, and 10 providers. 59 of those configurations report both an AA Index and a mean API cost. The values below are copied from that snapshot. AI Charts does not recalculate Artificial Analysis scores.

What this snapshot measures

AA Index is the snapshot's overall 0–100 score across code changes, terminal work, and repository understanding. Artificial Analysis also reports DeepSWE, Terminal-Bench v2, and SWE-Atlas-QnA as separate metrics. Those component scores are not combined here. This note uses AA Index because it is the composite the source publishes for the same configuration that also carries a task-level cost.

API cost is the mean API cost in US dollars for the evaluated task configuration. It is not a subscription price, a latency guarantee, or a production invoice. Active time and total token use exist in the same records and are left for the comparison chart.

Each row is a specific combination of model, agent harness, and effort setting. A model name without the harness and setting is an incomplete citation. Two rows that share a model and differ only in setting are different observations.

Highest AA Index configurations

The highest AA Index in this snapshot is 66.7 for Opus 5 on Claude Code at the xhigh setting, with a mean API cost of $8.23 per task.

Highest AA Index configurations in the Artificial Analysis snapshot retrieved Aug 18, 2026, 11:03 AM UTC
ModelAgentSettingAA IndexCost
Opus 5Claude Codexhigh66.7$8.23
GPT-5.6 SolCodexmax66.6$7.08
Fable 5 (with fallback)Claude Codemax65.8$11.70
Opus 5Claude Codemax65.5$8.95
GPT-5.6 SolCodexxhigh65.1$5.24
Grok 4.5Grok Buildhigh64.4$2.59
GPT-5.6 SolCodexhigh64.1$4.14
Opus 5Claude Codehigh63.4$3.80
GPT-5.6 TerraCodexmax62.3$2.21
Opus 5Claude Codemedium61.9$3.14

These are the highest stored AA Index scores, not a claim that the same systems lead on DeepSWE, Terminal-Bench v2, or SWE-Atlas-QnA. The dataset page lists the highest available score for each of those metrics separately and includes every configuration in the snapshot.

Cost and AA Index on the frontier

A higher AA Index usually comes with a higher mean task cost in this snapshot, but not every expensive configuration is on the useful edge. The cost/performance frontier keeps a configuration only when no other configuration is both cheaper and at least as strong on AA Index.

AA Index versus cost frontier in the Artificial Analysis snapshot retrieved Aug 18, 2026, 11:03 AM UTC
ModelAgentSettingAA IndexCost
GPT-5.6 LunaCodexlow25.1$0.042
Composer 2Cursor CLIdefault27.5$0.044
DeepSeek V4 FlashCodexmax55.5$0.071
GPT-5.6 LunaCodexmax58.7$0.313
GPT-5.6 TerraCodexmax62.3$2.21
Grok 4.5Grok Buildhigh64.4$2.59
GPT-5.6 SolCodexxhigh65.1$5.24
GPT-5.6 SolCodexmax66.6$7.08
Opus 5Claude Codexhigh66.7$8.23

The frontier in this snapshot has 9 configurations. The cheapest points are low-cost, lower-score runs. DeepSeek V4 Flash on Codex at the max setting is the first large AA Index increase that remains inexpensive. After that, each step buys a smaller AA Index gain at a higher mean task cost, ending at Opus 5 on Claude Code.

That sequence is AI Charts analysis of the stored pairs. Artificial Analysis does not publish a frontier ranking. The frontier can change when the next validated snapshot adds, removes, or reprices a configuration.

AA Index per dollar is a derived view

Dividing AA Index by mean API cost produces a derived ratio. It is not an Artificial Analysis metric. The ratio favors cheap configurations and can rank a low score above a stronger but more expensive run.

Highest derived AA Index per dollar in the Artificial Analysis snapshot retrieved Aug 18, 2026, 11:03 AM UTC
ModelAgentSettingAA IndexCostAA Index / $
DeepSeek V4 FlashCodexmax55.5$0.071784.7
Composer 2Cursor CLIdefault27.5$0.044629.6
GPT-5.6 LunaCodexlow25.1$0.042601.4
Composer 2.5Cursor CLIdefault38.2$0.082464.4
GPT-5.6 LunaCodexmedium42.4$0.095447.4
GPT-5.6 LunaCodexnone20.4$0.070290.1
GPT-5.6 LunaCodexhigh51.4$0.192268.2
GPT-5.6 LunaCodexxhigh54.7$0.251217.6
GPT-5.6 LunaCodexmax58.7$0.313187.2
DeepSeek V4 ProClaude Codehigh31.4$0.272115.6

Use the ratio only to find inexpensive configurations that still have a recorded AA Index. Use the frontier when the question is which configurations are not strictly worse on both cost and score.

How to use these numbers

Use this note when you need a sourced answer to a cost and quality question on the current coding-agent snapshot. Open the comparison chart to change axes, pin a model, or inspect provider ranges. Open the dataset page for provenance, benchmark definitions, and the full configuration table.

MirrorCode and SlopCodeBench answer different questions: project-scale reimplementation under large budgets, and quality change as an agent repeatedly extends its own code. They are not substitutes for the AA Index and cost pairs stored here.

Limits of the comparison

  • Artificial Analysis defines and operates the evaluations. AI Charts is an independent visualization and is not affiliated with Artificial Analysis or the listed providers.
  • Scores and costs belong to the named model, harness, setting, task set, and evaluation version on the retrieval date. They do not establish results for every repository or production workflow.
  • Mean task cost is not a price quote. Prompt mix, retry policy, caching, and live API prices can differ from the evaluation.
  • AA Index is a composite. A configuration can lead on the index and trail on a component benchmark.
  • This is a checked snapshot, not a live mirror. Cite the retrieval timestamp when quoting a value.

Sources

  1. Coding AgentsArtificial Analysis, 2026. The public coding-agents comparison is the upstream source of the checked AI Charts snapshot. Model names, agent harnesses, settings, AA Index scores, and mean API costs are Artificial Analysis measurements.

Results describe the named model, harness, task set, budget, and evaluation version. They do not establish performance on every production repository.