Source and refresh
The source is the public Artificial Analysis coding-agents comparison. AI Charts retrieved this snapshot on . The site checks for a new source snapshot daily. The displayed retrieval time changes only when a validated snapshot is stored.
The most recent retained model, variant, or material benchmark change was detected on . This meaningful-update time is separate from the daily retrieval check.
The checked dataset contains 59 model-agent configurations across 28 models, 9 agent harnesses, and 10 model providers. The chart and the JSON download use this same checked snapshot.
Benchmark definitions
Each benchmark is shown on the 0–100 scale stored in the snapshot. The metrics evaluate different tasks and should be interpreted separately.
AA Index
Overall performance across code changes, terminal work, and repository understanding.
DeepSWE
Long-horizon software engineering tasks scored with automated code verification.
Terminal-Bench v2
Agentic terminal-use tasks scored with automated test-suite verification.
SWE-Atlas-QnA
Repository-understanding questions scored with a strict resolve verifier.
Current leaders
These are the highest available scores in the retrieved snapshot, one row per benchmark. They are observations of the named model, agent harness, and effort setting rather than general model ranks.
| Benchmark | Model | Agent | Provider | Setting | Score |
|---|---|---|---|---|---|
| AA Index | Opus 5 | Claude Code | Anthropic | xhigh | 66.7 |
| DeepSWE | GPT-5.6 Sol | Codex | OpenAI | max | 68.7 |
| Terminal-Bench v2 | GPT-5.6 Sol | Codex | OpenAI | max | 87.7 |
| SWE-Atlas-QnA | Opus 5 | Claude Code | Anthropic | xhigh | 54.8 |
Normalization method
The refresh job reads the source page's public data payload, validates every source row, and maps it into a versioned owned schema. Benchmark reward proportions are represented as 0–100 scores. Mean task cost stays in US dollars, mean active wall time stays in seconds, and mean total token use stays as a token count.
Provider identifiers, model effort settings, stable series keys, and sort order are normalized for the chart. The refresh is rejected when duplicate records, major row loss, stable-key loss, or substantial metric-coverage regressions are detected. AI Charts does not recalculate the source benchmark outcomes.
Limitations
- Artificial Analysis defines and operates the upstream evaluations. AI Charts is an independent visualization and is not affiliated with Artificial Analysis or the listed providers.
- Scores depend on the named model, agent harness, effort setting, task set, and evaluation version. They do not establish results for every software repository or production workflow.
- Cost, duration, and token values are task-level means from the source evaluation. They are not price or latency guarantees.
- This is a daily checked snapshot, not a real-time mirror. Use the retrieval timestamp when citing a value.