Skip to dataset details

Coding-agent benchmark dataset

Download the current Artificial Analysis coding-agent benchmark snapshot, with benchmark definitions, provenance, leaders, methodology, and limitations.

Download JSON

Source and refresh

The source is the public Artificial Analysis coding-agents comparison. AI Charts retrieved this snapshot on . The site checks for a new source snapshot daily. The displayed retrieval time changes only when a validated snapshot is stored.

The most recent retained model, variant, or material benchmark change was detected on . This meaningful-update time is separate from the daily retrieval check.

The checked dataset contains 59 model-agent configurations across 28 models, 9 agent harnesses, and 10 model providers. The chart and the JSON download use this same checked snapshot.

Benchmark definitions

Each benchmark is shown on the 0–100 scale stored in the snapshot. The metrics evaluate different tasks and should be interpreted separately.

AA Index

Overall performance across code changes, terminal work, and repository understanding.

DeepSWE

Long-horizon software engineering tasks scored with automated code verification.

Terminal-Bench v2

Agentic terminal-use tasks scored with automated test-suite verification.

SWE-Atlas-QnA

Repository-understanding questions scored with a strict resolve verifier.

Current leaders

These are the highest available scores in the retrieved snapshot, one row per benchmark. They are observations of the named model, agent harness, and effort setting rather than general model ranks.

Highest score by benchmark in the current snapshot
BenchmarkModelAgentProviderSettingScore
AA IndexOpus 5Claude CodeAnthropicxhigh66.7
DeepSWEGPT-5.6 SolCodexOpenAImax68.7
Terminal-Bench v2GPT-5.6 SolCodexOpenAImax87.7
SWE-Atlas-QnAOpus 5Claude CodeAnthropicxhigh54.8

Normalization method

The refresh job reads the source page's public data payload, validates every source row, and maps it into a versioned owned schema. Benchmark reward proportions are represented as 0–100 scores. Mean task cost stays in US dollars, mean active wall time stays in seconds, and mean total token use stays as a token count.

Provider identifiers, model effort settings, stable series keys, and sort order are normalized for the chart. The refresh is rejected when duplicate records, major row loss, stable-key loss, or substantial metric-coverage regressions are detected. AI Charts does not recalculate the source benchmark outcomes.

Limitations

  • Artificial Analysis defines and operates the upstream evaluations. AI Charts is an independent visualization and is not affiliated with Artificial Analysis or the listed providers.
  • Scores depend on the named model, agent harness, effort setting, task set, and evaluation version. They do not establish results for every software repository or production workflow.
  • Cost, duration, and token values are task-level means from the source evaluation. They are not price or latency guarantees.
  • This is a daily checked snapshot, not a real-time mirror. Use the retrieval timestamp when citing a value.
  1. Download the current JSON snapshotVersioned records, provenance, retrieval time, and bounded update history used by the production chart.
  2. Artificial Analysis coding-agents sourceThe upstream comparison from which the checked snapshot is derived.
  3. Refresh and normalization source codeThe public parser, normalization rules, validation guards, and update-detection logic.