Skip to notes
AI Charts
Theme
Appearance

Introducing AI Charts

AI Charts plots published AI benchmark results, and each benchmark snapshot records its source, named version, and retrieval date.

Drafted with AI from the source code and reviewed by Claude Opus 5.5 (claude-opus-5-5) editorial review.

Seven small ivory forms, from a sphere to a pyramid, sit on separate charcoal plinths along one shelf, each with a blank tag hanging from a thin brass rail.
Each result on AI Charts keeps its own source, version, and retrieval date, and each benchmark stays on its own scale. AI Charts editorial illustration · Slopcamera with GPT Image 2

AI Charts plots published benchmark results for AI models and coding agents against cost, time, and token use. Each chart names the source of its numbers and the date they were retrieved, because a score is hard to use without knowing what produced it and when.

A score means little without its setup

The same model can appear on a coding leaderboard several times with different results. That is because a coding-agent result belongs to a whole configuration: the model, the agent harness that runs it, and the effort setting. The benchmark also has a version, each run has a cost, and the table was true on one date. A leaderboard screenshot usually loses most of that.

On the coding-agent chart, each point is one named configuration. Hover over it to see the model, harness, and effort setting, the scores, and the cost, time, and total token use where the source reports them. In the benchmarks library, results also show the uncertainty interval the source reports, labeled with the source's own interval type.

Who it is for

AI Charts is for someone choosing a model or coding agent who wants to weigh score against cost, time, or tokens. A typical question is which configurations score about as well as the leader for much less per task. The chart answers it with the cost frontier: a configuration is on the frontier only when nothing cheaper scores at least as well. The note Highest AA Index and lowest cost pick different coding agents walks through that trade-off on one dated snapshot.

If you want one overall rank across every kind of task, use something else, because AI Charts builds no composite score of its own. Where a source publishes an index, such as Artificial Analysis's, AI Charts shows that index as the source defines it. Reasoning, research, memory, image, video, and audio results stay on their own scales. AI Charts also cannot tell you how a model will do on your own codebase. The note Why a coding-agent high score still needs a holdout explains why a test set the model never saw is still worth building.

What is on the site today

The homepage chart plots model configurations by Artificial Analysis Intelligence Index score against cost or output tokens per task. In the snapshot retrieved Sep 23, 2026, it uses Intelligence Index v4.3.2, and the earlier v4.1.1 results stay a separate dataset. Scores are never relabeled from one version to another.

The coding chart plots results from the Artificial Analysis Coding Agent Index v1.5 and its three components, DeepSWE v1.1, Terminal-Bench 4, and SWE-Atlas-QnA, against cost, duration, or total token use. Pin a model to see the configurations that score near it, or pin a provider to see its range. The site's coding standard is the official, version-pinned Terminal-Bench 4.0 snapshot from the benchmark's owners, which sits in the benchmarks library as its own cohort. Artificial Analysis runs Terminal-Bench 4 in its own harness, so its scores and the owners' results are kept apart.

The benchmarks library covers reasoning, research, memory, images, video, audio, and world models. It labels each entry as a charted result, a source guide, or an emerging evaluation. A source guide describes a benchmark whose scores AI Charts has not imported, and it shows no numbers.

The data page lists, for every entry, the question it answers, what it measures, the source, the version, which comparisons are valid, and the limits. Charted entries link a JSON download of the plotted data.

The notes each take one benchmark, study, or result and explain what it measures and how far the evidence goes. What HarnessTax’s same-model cost gap measures covers a study that ran the same models in different harnesses on two public suites. What Terminal-Bench-Science’s 30% result measures reads a science benchmark's top result against its cost and token use. Each note cites its primary sources and names the configuration and date behind the results it discusses.

Where the catalog is going

The aim is a catalog in which any published benchmark result someone might use to pick a model can be read with its configuration, version, cost, and date. Benchmarks that have only a source guide today are meant to become charts once their data can be checked the same way. New notes will follow the questions readers bring to the charts, and each benchmark will keep its own scale.

What AI Charts does not do, and its status

AI Charts does not run evaluations. The scores, costs, and token counts come from the benchmark owners and aggregators it cites, and it is not affiliated with them or with the model providers in the data. Cost figures keep the source's denominator, such as per task or per full evaluation, so two costs are comparable only when that denominator matches. The charts show dated snapshots: the Intelligence Index snapshot is checked for updates every four hours and the coding-agent snapshot daily, and the earlier v4.1.1 data is frozen. Some vendor-run results, such as CursorBench, appear as supplemental evidence for a model running inside that vendor's product, not as an independent standard.

The benchmark charts and notes are live at aicharts.io. AI Charts also includes a local tool that measures your own coding agents' token use. Its status is In development: build it from source, since there is no packaged release yet.

Sources

  1. Coding AgentsArtificial Analysis, 2026. The public coding-agents comparison is the source of the AI Charts coding-agent snapshot. Model names, agent harnesses, settings, AA Index scores, and mean API costs are Artificial Analysis measurements.
  2. LLM LeaderboardArtificial Analysis, 2026. The public models leaderboard is the source of the AI Charts Intelligence Index snapshot. Scores, per-task costs, and output tokens are Artificial Analysis measurements under Intelligence Index v4.3.2.
  3. Terminal-BenchHarbor Framework, 2026. The benchmark owners’ repository publishes Terminal-Bench and its versioned task releases. The version-pinned Terminal-Bench 4.0 cohort in the AI Charts benchmarks library comes from the owners, separately from Artificial Analysis’s own Terminal-Bench 4 runs.

Figures come from the cited primary sources and the AI Charts datasets. AI Charts did not rerun the reported benchmarks. Reported results apply to the named source, workload, configuration, and observation date. They do not establish performance on every task or product.