Skip to notes
AI Charts
Theme
Appearance

AI model and agent benchmark analysis

Sourced analysis of AI model and agent benchmarks: what each evaluation measures, what the results show, and where the evidence stops. The first collection focuses on coding agents.

Explore the coding-agent chart

Articles

16 sourced analysis articles

Seven small ivory forms, from a sphere to a pyramid, sit on separate charcoal plinths along one shelf, each with a blank tag hanging from a thin brass rail.

Introducing AI Charts

AI Charts plots published AI benchmark results, and each benchmark snapshot records its source, named version, and retrieval date.

A matte ivory sphere rests on a charcoal ledge above two separate dark pools, each holding its own reflection of it.

GPT-6 Sol scores 56.7 on the coding-agent chart at $2.99 a task

Codex · GPT-6 Sol (max) scores 56.7 on the coding-agent AA Index at $2.99 per task, seventh of 20 configurations, and GPT-6 Sol (max) scores 47.5 on the Intelligence Index at $1.06 per task. Each row sits on its own chart’s cost frontier.

A large matte ivory stone rests across two separate charcoal plinths of different heights, each edged by its own short brass line.

What Grok 4.7’s 56 on the coding-agent chart measures

On the September 25, 2026 snapshots, Grok Build · Grok 4.7 (xhigh) scores 56 on the coding-agent AA Index and Grok 4.7 (xhigh) scores 46 on the Intelligence Index. Different harnesses, task sets, and costs sit behind the two numbers, and cheaper configurations score higher on both charts.

A small ivory sphere rests on a low charcoal step beside three tall dark pillars, on a staircase traced by one thin brass line.

What MiMo-V2.6-Pro’s 46 at $0.13 per task measures

On the September 22, 2026 Intelligence Index snapshot, Xiaomi’s open-weights flagship scores within a point of GPT-5.6 Sol and Grok 4.7 at less than a tenth of their cost per task. Its cybersecurity lead is on one kind of task, measured by Xiaomi.

A thick charcoal band carries a pale inner path that curves and meets a thinner side path, then continues as one quieter channel.

What Devin Fusion’s 39% saving measures

Cognition’s 39% is one comparison: an Astra-led Fusion pair against Codex on the Artificial Analysis index, at a lower score. The same post’s other reported savings run from 11% to 46%.

Method

Sources

Each note starts with primary or first-party evidence. Material claims link directly to those sources.

Changing results

Leaderboard values are paired with their observation date and named configuration. They can change after publication.

Limits

Methodology limits and interpretation are kept near the results they qualify. Benchmark performance is not treated as a general production claim.