Skip to dataset details
aicharts
Theme
Appearance

Benchmark data and method

Versioned AI benchmark charts and source guides across coding, reasoning, research, memory, images, video, audio, and world models. Each measured cohort retains its source, configuration, score unit, and comparison limits.

Benchmark atlas: charts and source guides

The atlas covers 62 benchmarks and 29 charted evaluation cohorts. Each chart uses one source, score definition, and version. A source guide explains an evaluation whose results are not yet charted here. Historical research cohorts remain labeled; retrieving a paper today does not mean its models were evaluated today.

Download the catalog JSON to discover the available benchmark IDs, comparison rules, source dates, and individual dataset URLs. Each dataset download includes the same observations, units, costs, uncertainty labels, and configurations used by its chart. This is a checked snapshot distribution, not a live model API.

Cite the benchmark owner and source version when quoting measurements. aicharts publishes these chart projections; third-party measurements retain their source terms. The software license does not grant a new license to third-party data.

Terminal-Bench 4 · 4.0.0 · 13 charted results

Which agent can finish difficult terminal work? Software, systems, CAD, scientific computing, and formal proof tasks executed in a terminal.

Measure
Task success over 66 tasks and five trials per configuration, with source-reported 95% confidence intervals.
Comparison rule
aicharts’ primary terminal benchmark. Compare exact version 4.0.0 with the named model, agent version, and effort.
Evaluation cohort
66 tasks × 5 trials per configuration. Model, agent, and effort all affect results. The owner reports 95% confidence intervals; its interval method and price basis are unspecified.
Score
Task success (%); higher is better.
Source retrieved
Source revision
c7878bd7cbd05d1dd21586b83eb3a1929a412312
Cost basis
USD for the full 330-trial evaluation
  • An agent and model are evaluated together; this is not a model-only ranking.
  • Different harnesses and effort settings remain separate configurations.
  • Cost covers the full 330-trial evaluation. The source does not specify its pricing or confidence-interval method.

Explore Terminal-Bench 4 · Harbor Framework · Methodology · Download this dataset

Artificial Analysis Intelligence Index · 4.3.2 · 101 charted results

How much capability do I get for the cost? Broad model capability alongside the cost and output tokens used to achieve it, within the same evaluation cohort.

Measure
The publisher’s v4.3.2 index across ten evaluations: agents 30%, coding 20%, scientific reasoning 20%, and general capability 30%.
Comparison rule
Compare configurations only within v4.3.2. Its evaluation roster and weights differ from v4.1.1; the historical scores are not on a continuous scale with these scores.
Evaluation cohort
Ten-evaluation v4.3.2 cohort. Compare the same reasoning configuration and native per-task measurements. Historical v4.1.1 scores remain separate; a later index revision requires a new versioned cohort.
Score
Intelligence Index (index points); higher is better.
Source retrieved
Cost basis
USD per Intelligence Index task
  • An index reflects the publisher’s task mix and weights, not every use case. Reasoning effort changes the evaluated configuration.
  • Both resources are publisher-reported per-task measures. Output tokens include answer and reasoning; cost also includes input and cache traffic.
  • Only current, non-estimated configurations with complete native measurements are retained. Missing positive cost is not free. No per-configuration uncertainty is supplied.
  • The Terminal-Bench component uses Artificial Analysis’s own evaluation harness and must not be merged with the Harbor submission leaderboard.

Explore Artificial Analysis Intelligence Index · Artificial Analysis · Methodology · Download this dataset

Image generation · Arena · Overall · Bradley–Terry · 80 charted results

Which image generators lead Arena’s preference ratings? Compare published preference ratings for images generated from text, with uncertainty and vote counts.

Measure
Published Bradley–Terry preference rating; higher is better. Source confidence intervals and vote counts included.
Comparison rule
Text-to-image, overall category only. One complete, dated publisher cohort; no blending with other Arena tracks or diagnostic benchmarks.
Evaluation cohort
Text-to-image overall ratings, published 2026-09-24. Compare within this dated pool; retain settings and confidence intervals. Preliminary and AutoEval flags are not provided. Arena (lmarena-ai), Arena leaderboard dataset, CC-BY-4.0 (https://creativecommons.org/licenses/by/4.0/). Selected overall media cohorts and projected fields for AI Charts; published ratings, intervals, and vote counts unchanged.
Score
Arena preference rating (Arena points); higher is better.
Source retrieved
Source observation date
Source revision
6931a5ef0714c8d4edf09d79998fef6674ed4e09
  • Ratings depend on the opponent pool and prompt mix. Compare only within this track and publication date; ratings are not percentages.
  • Overlapping confidence intervals do not establish a clear winner. Vote count is per model, not a count of unique people or prompts.
  • The licensed export omits preliminary and AutoEval status flags. A row’s evidence status cannot be inferred from its vote count.
  • This extract has no matched price or latency data. Keep audio, resolution, effort, web-search, and agent settings in the model labels.
  • Arena supports AutoEval proxy votes for image generation; the export does not identify affected rows. Preference does not guarantee correct composition or faithful editing.

Explore Image generation · Arena · Arena · Methodology · Download this dataset

Image editing · Arena · Overall · Bradley–Terry · 56 charted results

Which image editors lead Arena’s preference ratings? A separate preference comparison for editing an existing image, preserving each published model setting.

Measure
Published Bradley–Terry preference rating; higher is better. Source confidence intervals and vote counts included.
Comparison rule
Image editing, overall category only. One complete, dated publisher cohort; no blending with other Arena tracks or diagnostic benchmarks.
Evaluation cohort
Image editing overall ratings, published 2026-09-22. Compare within this dated pool; retain settings and confidence intervals. Preliminary and AutoEval flags are not provided. Arena (lmarena-ai), Arena leaderboard dataset, CC-BY-4.0 (https://creativecommons.org/licenses/by/4.0/). Selected overall media cohorts and projected fields for AI Charts; published ratings, intervals, and vote counts unchanged.
Score
Arena preference rating (Arena points); higher is better.
Source retrieved
Source observation date
Source revision
6931a5ef0714c8d4edf09d79998fef6674ed4e09
  • Ratings depend on the opponent pool and prompt mix. Compare only within this track and publication date; ratings are not percentages.
  • Overlapping confidence intervals do not establish a clear winner. Vote count is per model, not a count of unique people or prompts.
  • The licensed export omits preliminary and AutoEval status flags. A row’s evidence status cannot be inferred from its vote count.
  • This extract has no matched price or latency data. Keep audio, resolution, effort, web-search, and agent settings in the model labels.
  • Arena supports AutoEval proxy votes for image generation; the export does not identify affected rows. Preference does not guarantee correct composition or faithful editing.

Explore Image editing · Arena · Arena · Methodology · Download this dataset

Video generation · Arena · Overall · Bradley–Terry · 48 charted results

Which text-to-video systems lead Arena’s preference ratings? Compare video generators in the same text-prompt preference pool, keeping audio, resolution, and agent labels intact.

Measure
Published Bradley–Terry preference rating; higher is better. Source confidence intervals and vote counts included.
Comparison rule
Text-to-video, overall category only. One complete, dated publisher cohort; no blending with other Arena tracks or diagnostic benchmarks.
Evaluation cohort
Text-to-video overall ratings, published 2026-09-22. Compare within this dated pool; retain settings and confidence intervals. Preliminary and AutoEval flags are not provided. Arena (lmarena-ai), Arena leaderboard dataset, CC-BY-4.0 (https://creativecommons.org/licenses/by/4.0/). Selected overall media cohorts and projected fields for AI Charts; published ratings, intervals, and vote counts unchanged.
Score
Arena preference rating (Arena points); higher is better.
Source retrieved
Source observation date
Source revision
6931a5ef0714c8d4edf09d79998fef6674ed4e09
  • Ratings depend on the opponent pool and prompt mix. Compare only within this track and publication date; ratings are not percentages.
  • Overlapping confidence intervals do not establish a clear winner. Vote count is per model, not a count of unique people or prompts.
  • The licensed export omits preliminary and AutoEval status flags. A row’s evidence status cannot be inferred from its vote count.
  • This extract has no matched price or latency data. Keep audio, resolution, effort, web-search, and agent settings in the model labels.
  • Video preference does not measure physical plausibility, interactive world simulation, or camera-control accuracy.

Explore Video generation · Arena · Arena · Methodology · Download this dataset

Image-to-video · Arena · Overall · Bradley–Terry · 48 charted results

Which systems turn an image into a preferred video? An image-conditioned video preference comparison, separate from generation that starts with a text prompt alone.

Measure
Published Bradley–Terry preference rating; higher is better. Source confidence intervals and vote counts included.
Comparison rule
Image-to-video, overall category only. One complete, dated publisher cohort; no blending with other Arena tracks or diagnostic benchmarks.
Evaluation cohort
Image-to-video overall ratings, published 2026-09-22. Compare within this dated pool; retain settings and confidence intervals. Preliminary and AutoEval flags are not provided. Arena (lmarena-ai), Arena leaderboard dataset, CC-BY-4.0 (https://creativecommons.org/licenses/by/4.0/). Selected overall media cohorts and projected fields for AI Charts; published ratings, intervals, and vote counts unchanged.
Score
Arena preference rating (Arena points); higher is better.
Source retrieved
Source observation date
Source revision
6931a5ef0714c8d4edf09d79998fef6674ed4e09
  • Ratings depend on the opponent pool and prompt mix. Compare only within this track and publication date; ratings are not percentages.
  • Overlapping confidence intervals do not establish a clear winner. Vote count is per model, not a count of unique people or prompts.
  • The licensed export omits preliminary and AutoEval status flags. A row’s evidence status cannot be inferred from its vote count.
  • This extract has no matched price or latency data. Keep audio, resolution, effort, web-search, and agent settings in the model labels.
  • Video preference does not measure physical plausibility, interactive world simulation, or camera-control accuracy.

Explore Image-to-video · Arena · Arena · Methodology · Download this dataset

Artificial Analysis Intelligence Index · 4.1.1 · 135 charted results

How much capability do I get for the cost? A versioned snapshot of broad model capability, with the output tokens and cost used to achieve it.

Measure
An owner-defined index across nine evaluations of agents, coding, scientific reasoning, and general knowledge.
Comparison rule
Compare models within the retained v4.1.1 cohort. Later index versions use different evaluations or weights and require a separate chart.
Evaluation cohort
Historical v4.1.1 snapshot of the nine-evaluation index. The publisher introduced v4.3 on September 7; those results are not merged here. Effort settings and missing costs remain explicit.
Score
Intelligence Index (index points); higher is better.
Source retrieved
Cost basis
USD per Intelligence Index task
  • The publisher introduced v4.3 on September 7, 2026. This historical v4.1.1 chart is not a current-version ranking.
  • The index reflects the publisher’s task mix and weights, not every use case.
  • Reasoning effort changes the system being evaluated. Missing cost does not mean free.

Explore Artificial Analysis Intelligence Index · Artificial Analysis · Methodology · Download this dataset

Terminal-Bench-Science · 0.1.0 · 17 charted results

Can an agent carry out scientific work? Research tasks that require scientific software, computation, and evidence across five domains.

Measure
Resolution rate across 70 tasks and three trials, with separate domain results.
Comparison rule
Compare the exact 0.1.0 cohort. Error bars show one standard error, which is narrower than a 95% confidence interval.
Evaluation cohort
70 science tasks × 3 trials. Error bars show one binomial standard error, not a 95% confidence interval. Model and agent versions, price basis, and several run limits are unspecified.
Score
Resolution rate (%); higher is better.
Source retrieved
Source observation date
Source revision
f81afac4f11048e77a15dfc8fb1dbfb897fea0ce
Cost basis
USD for the full 210-trial evaluation
  • Model and agent versions and several run limits are not specified by the source.
  • Cost covers the full 210-trial evaluation; source aggregate and domain costs can differ.
  • A benchmark success does not establish the scientific validity of arbitrary new research.

Explore Terminal-Bench-Science · Terminal-Bench-Science · Methodology · Download this dataset

DeepSWE v1.1 · AA evaluation · 1.1 · 20 charted results

Can an agent complete long software changes? Long-horizon software engineering, measured inside Artificial Analysis’s coding-agent evaluation.

Measure
Success on software engineering tasks with automated code verification.
Comparison rule
Keep this Artificial Analysis cohort separate from DataCurve’s direct mini-swe-agent leaderboard and other harnesses.
Evaluation cohort
Each point is an Artificial Analysis coding-agent configuration. Cost covers the source’s coding suite, not an isolated run of the selected benchmark. Component results stay separate from standalone cohorts published by other evaluators.
Score
DeepSWE v1.1 (%); higher is better.
Source retrieved
Cost basis
USD per task across the AA coding suite
  • Cost and time describe the AA coding suite, not the isolated DeepSWE task set.
  • The checked source does not report per-configuration uncertainty.

Explore DeepSWE v1.1 · AA evaluation · Artificial Analysis · Download this dataset

Artificial Analysis Coding Agent Index · 1.5 · 20 charted results

Which coding configuration covers a range of tasks? The publisher’s combined view of software changes, terminal work, and repository understanding.

Measure
The source’s Coding Agent Index alongside mean evaluation cost, time, and token use.
Comparison rule
Read the source-defined index within this coding-suite snapshot. It is distinct from the general Intelligence Index.
Evaluation cohort
Each point is an Artificial Analysis coding-agent configuration. Cost covers the source’s coding suite, not an isolated run of the selected benchmark. Component results stay separate from standalone cohorts published by other evaluators.
Score
Coding Agent Index (index points); higher is better.
Source retrieved
Cost basis
USD per task across the AA coding suite
  • The combined score reflects the evaluator’s task mix.
  • Different agent harnesses and effort settings affect both performance and resource use.
  • Component scores remain specific to this Artificial Analysis cohort.

Explore Artificial Analysis Coding Agent Index · Artificial Analysis · Download this dataset

Terminal-Bench 4 · AA evaluation · 4.0 · 20 charted results

How did these coding configurations perform on terminal tasks? Terminal-Bench 4 results retained inside the Artificial Analysis coding-suite comparison.

Measure
Task success on Terminal-Bench 4 with the named agent and model.
Comparison rule
Keep this Artificial Analysis cohort separate from the standalone owner-published Terminal-Bench 4 cohort.
Evaluation cohort
Each point is an Artificial Analysis coding-agent configuration. Cost covers the source’s coding suite, not an isolated run of the selected benchmark. Component results stay separate from standalone cohorts published by other evaluators.
Score
Terminal-Bench 4 (%); higher is better.
Source retrieved
Cost basis
USD per task across the AA coding suite
  • Cost and time cover the AA coding suite, not this isolated benchmark.
  • The checked source does not report per-configuration uncertainty.

Explore Terminal-Bench 4 · AA evaluation · Artificial Analysis · Download this dataset

SWE-Atlas-QnA · swe-atlas-qna · 20 charted results

Can an agent understand a software repository? Repository questions evaluated with a strict answer verifier.

Measure
Successful answers to repository-understanding questions in the AA evaluation.
Comparison rule
Compare the retained SWE-Atlas-QnA source cohort and exact agent configuration.
Evaluation cohort
Each point is an Artificial Analysis coding-agent configuration. Cost covers the source’s coding suite, not an isolated run of the selected benchmark. Component results stay separate from standalone cohorts published by other evaluators.
Score
SWE-Atlas-QnA (%); higher is better.
Source retrieved
Cost basis
USD per task across the AA coding suite
  • Repository understanding is narrower than successfully changing and shipping software.
  • The source does not expose a finer immutable dataset revision in this snapshot.
  • Cost and time cover the AA coding suite, not this isolated benchmark.

Explore SWE-Atlas-QnA · Artificial Analysis · Download this dataset

Vals Index · Vals AI · 2 · private test set · 66 charted results

Which model does the most economically valuable work? A GDP-weighted average of agentic performance across finance, coding, and legal tasks.

Measure
The publisher's composite of seven evaluations, weighted by each sector's share of U.S. GDP: finance 8.0, coding 5.6, and legal 1.2, over a denominator of 14.8.
Comparison rule
Compare within index version 2. Versions 1 and 2 use different component benchmarks, so their scores are not a series.
Evaluation cohort
One publisher-defined composite. Cost is the average dollar cost of one test across the weighted components, not the cost of a task you would run. Model names are the identifiers Vals publishes, not aicharts canonical names.
Score
Accuracy (%); higher is better.
Source retrieved
Source observation date
Source revision
vals_index v2
Cost basis
Average cost per test (USD)
  • This is the publisher's composite, not an aicharts ranking. The sector weights are a deliberate simplification of how AI reaches the economy.
  • The Code Migration component scores a fixed 60-task subset of the published 120-task run, not the full benchmark.
  • Five of the seven components are private evaluations that no independent party can reproduce.
  • Model names are the identifiers Vals publishes, not aicharts canonical names.

Explore Vals Index · Vals AI · Vals AI · Methodology · Download this dataset

Finance Agent (v2) · Vals AI · 2 · private test set · 70 charted results

Can an agent do a financial analyst's work? Multi-step analyst tasks covering modeling, comparables, earnings, and disclosure reading.

Measure
Severity-weighted partial credit across ten task families in the FAB v2 harness, averaged over repeated runs.
Comparison rule
Compare within Finance Agent v2. The v1 board used a different task set and is not a continuation of this one.
Evaluation cohort
Private evaluation run by Vals. Dollar values are the publisher's average cost per test under its own harness. Model names are the identifiers Vals publishes, not aicharts canonical names.
Score
Accuracy (%); higher is better.
Source retrieved
Source observation date
Source revision
finance_agent v2
Cost basis
Average cost per test (USD)
  • The test set is private, so no independent party can reproduce these scores.
  • Task-family scores in the inspector are components of the headline number, not separate rankings.
  • Model names are the identifiers Vals publishes, not aicharts canonical names.

Explore Finance Agent (v2) · Vals AI · Vals AI · Methodology · Download this dataset

Tax Agent Bench · Vals AI · 1 · private test set · 61 charted results

Can an agent answer a research-grade tax question? US corporate tax questions covering fact patterns, rule lookup, calculations, forms, and controversy.

Measure
Accuracy across six tax task families, with a stricter all-pass variant reported alongside the headline score.
Comparison rule
Compare within Tax Agent Bench version 1. This board covers 22 configurations, fewer than the other Vals boards.
Evaluation cohort
Private evaluation run by Vals. Dollar values are the publisher's average cost per test. Model names are the identifiers Vals publishes, not aicharts canonical names.
Score
Accuracy (%); higher is better.
Source retrieved
Source observation date
Source revision
tax_agent_bench v1
Cost basis
Average cost per test (USD)
  • The test set is private, so no independent party can reproduce these scores.
  • US corporate tax only, and tax rules change with the filing year.
  • A score here is not tax advice and does not establish that an answer is filing-ready.
  • Model names are the identifiers Vals publishes, not aicharts canonical names.

Explore Tax Agent Bench · Vals AI · Vals AI · Methodology · Download this dataset

MedCode · Vals AI · 1 · private test set · 102 charted results

Can a model assign the right medical billing code? Medical coding for the billing process, scored one-shot rather than as an agent.

Measure
Accuracy on medical billing code assignment in a single-response setting.
Comparison rule
One-shot only. These scores are not comparable with the agentic Vals boards, which let a model take many steps.
Evaluation cohort
Private one-shot evaluation run by Vals. The publisher does not present a comparable cost for this board. Model names are the identifiers Vals publishes, not aicharts canonical names.
Score
Accuracy (%); higher is better.
Source retrieved
Source observation date
Source revision
medcode v1
  • Vals does not present a comparable cost for this board, so no cost axis is offered.
  • The board retains configurations from 2025 alongside current models, so the range spans more than one model generation.
  • The test set is private, so no independent party can reproduce these scores.
  • Model names are the identifiers Vals publishes, not aicharts canonical names.

Explore MedCode · Vals AI · Vals AI · Methodology · Download this dataset

ARC-AGI-2 · 2 · semi-private · 28 charted results

Can it solve an unfamiliar visual puzzle? Infer a rule from a few examples, then apply it to a new grid.

Measure
Percentage of novel abstract grid tasks solved in ARC Prize's semi-private evaluation.
Comparison rule
Compare the same dataset split and reasoning effort. This view selects seven recent model families.
Evaluation cohort
Semi-private set. Seven selected recent model families; cost is unpublished for these configurations.
Score
Tasks solved (%); higher is better.
Source retrieved
Source observation date
Source revision
SHA-256 b5cb5ec6e8c7547c4f11a4e61e5ed7e0f28a26208f87bd16c83d1703777a2627
  • Strong puzzle performance does not establish general intelligence.
  • Selected configurations have no published cost in this source.
  • Scores near the ceiling leave less room to distinguish leading systems.

Explore ARC-AGI-2 · ARC Prize · Methodology · Download this dataset

ARC-AGI-3 · Standard · 3 · semi-private · Standard · 12 charted results

Can it learn the rules by interacting? Explore an unfamiliar environment and solve it through a shared, minimal agent interface.

Measure
Action-efficiency score relative to the benchmark's human baseline, expressed as a percentage.
Comparison rule
Standard harness only. Compare the selected Astra, Sol, and Opus 5 configurations within this view.
Evaluation cohort
Standard harness only. Dollar values cover the full evaluation. Harness cohorts are shown separately.
Score
Action-efficiency score (%); higher is better.
Source retrieved
Source observation date
Source revision
SHA-256 d743d7a731ba6b9f102d3624d5fe82011630b90da4200ceb34bd56cf2b221e1c
Cost basis
Total evaluation cost (USD)
  • Cost is the full evaluation spend, not the cost of one task.
  • These are deterministic puzzle environments; they do not measure open-ended real-world competence.
  • The Provider Adapter harness is a separate comparison.

Explore ARC-AGI-3 · Standard · ARC Prize · Methodology · Download this dataset

ARC-AGI-3 · Provider Adapter · 3 · semi-private · Provider Adapter · 6 charted results

How much does native context management help? Astra's reasoning settings with its provider's conversation and compaction features enabled.

Measure
ARC-AGI-3 action-efficiency score using the Provider Adapter harness.
Comparison rule
Compare Astra's six effort settings here. These scores are not a ranking against Standard-harness runs.
Evaluation cohort
Provider Adapter harness only. Dollar values cover the full evaluation. Harness cohorts are shown separately.
Score
Action-efficiency score (%); higher is better.
Source retrieved
Source observation date
Source revision
SHA-256 d743d7a731ba6b9f102d3624d5fe82011630b90da4200ceb34bd56cf2b221e1c
Cost basis
Total evaluation cost (USD)
  • Only Astra is included in this selected adapter cohort.
  • Cost is total evaluation spend.
  • Near-perfect performance on this bounded test is not proof of general intelligence.

Explore ARC-AGI-3 · Provider Adapter · ARC Prize · Methodology · Download this dataset

DeepResearch Bench II · II · 132 tasks · 9 charted results

Which research agent produces a useful report? Compare information gathering, analysis, and presentation against expert-written research rubrics.

Measure
Owner-reported weighted rubric score across 132 tasks and 9,430 criteria; component scores appear in the inspector.
Comparison rule
Keep the exact research product and model generation. An o3-era research run does not represent today's OpenAI product.
Evaluation cohort
Nine selected systems. The table includes older products; run dates, exact model versions, costs, and confidence intervals are not consistently supplied.
Score
Weighted rubric score (%); higher is better.
Source retrieved
Source revision
SHA-256 ba0ccb0a89ff6c8f5d3c97861b072cdb8459c27fdaef2db90e725537f2869562
  • The source includes historical products and does not date every run.
  • A model judge evaluates reports; no intervals or comparable costs are published in this table.
  • This is a selected set of nine systems, not the full source leaderboard.

Explore DeepResearch Bench II · USTC / Metastone · Methodology · Download this dataset

LongMemEval-V2 · Small · 2 · small · paper baselines · 6 charted results

Can an agent reuse what it learned before? Compare memory systems that turn previous web-agent work into useful evidence for a later question.

Measure
Answer accuracy with a fixed reader. Query latency is shown for each memory method.
Comparison rule
Compare memory methods within the same history tier and reader configuration. Small and Medium are separate sets.
Evaluation cohort
Small history tier only. This compares six memory methods, not six foundation models. Latency is in the inspector; dollar cost is not published.
Score
Answer accuracy (%); higher is better.
Source retrieved
Source revision
Paper baselines; evaluation code 2cc8c540bdb87fe6761629b585e727e1c4704520
  • These are six paper baselines; the public submission leaderboard is still empty.
  • This measures memory-system plus reader performance, not a general model ranking.
  • Query latency excludes the wider application experience; no comparable dollar costs are published.

Explore LongMemEval-V2 · Small · UCLA LongMemEval team · Methodology · Download this dataset

LongMemEval-V2 · Medium · 2 · medium · paper baselines · 6 charted results

Can an agent reuse what it learned before? Compare memory systems that turn previous web-agent work into useful evidence for a later question.

Measure
Answer accuracy with a fixed reader. Query latency is shown for each memory method.
Comparison rule
Compare memory methods within the same history tier and reader configuration. Small and Medium are separate sets.
Evaluation cohort
Medium history tier only. This compares six memory methods, not six foundation models. Latency is in the inspector; dollar cost is not published.
Score
Answer accuracy (%); higher is better.
Source retrieved
Source revision
Paper baselines; evaluation code 2cc8c540bdb87fe6761629b585e727e1c4704520
  • These are six paper baselines; the public submission leaderboard is still empty.
  • This measures memory-system plus reader performance, not a general model ranking.
  • Query latency excludes the wider application experience; no comparable dollar costs are published.

Explore LongMemEval-V2 · Medium · UCLA LongMemEval team · Methodology · Download this dataset

LongMemEval · 1 · S / M · Source guide

Will it remember changing facts across conversations? Tests recall, updates, time, reasoning across sessions, and knowing when an answer is absent.

Measure
Answer accuracy on 500 questions across long conversation histories.
Comparison rule
Keep S and M separate; match reader, judge, retrieval budget, and history construction.
  • Retrieval recall@k is not answer accuracy.
  • Vendor memory-system scores often use different readers and judges.

Explore LongMemEval · LongMemEval authors · Methodology

LoCoMo · ACL 2024 · public 10-conversation set · Source guide

Can it recall details from a long-running conversation? A widely cited test for conversational memory, event summaries, and dialogue continuity.

Measure
Question-answering and conversation tasks with long, multi-session histories.
Comparison rule
Name the public subset, included question categories, scoring method, reader, and retrieval budget.
  • F1, model-judged correctness, and retrieval recall are different scores.
  • The small conversation set and differing category exclusions limit cross-report comparisons.

Explore LoCoMo · Snap Research / UNC · Methodology

LongBench v2 · 2 · 503 questions · Source guide

Can it reason over a very long document? Reading comprehension across documents, conversations, code repositories, and structured data.

Measure
Multiple-choice accuracy, with short, medium, and long context breakdowns.
Comparison rule
Match chain-of-thought setting, input truncation, length bucket, and model context limit.
  • Long-context reading does not measure persistent memory between sessions.
  • Advertised context capacity does not guarantee useful recall at that length.

Explore LongBench v2 · LongBench team · Methodology

BrowseComp · 2025 · 1,266 questions · Source guide

Can it find a hard-to-locate fact on the web? Persistent search and multi-step browsing, scored through short factual answers.

Measure
Accuracy on web questions that require extensive information seeking.
Comparison rule
Match search tools, context management, agent count, and retry policy; label developer-reported scores.
  • It does not directly evaluate the quality of a long research report.
  • Live search results and different harnesses make cross-release scores difficult to compare.

Explore BrowseComp · OpenAI · Methodology

BrowseComp-Plus · ACL 2026 · fixed corpus · Source guide

How good is the research agent when search access is controlled? A reproducible information-retrieval setting for difficult research questions.

Measure
Answer accuracy and retrieval effectiveness over a released document collection.
Comparison rule
Keep corpus revision, retriever, retrieval budget, and agent configuration fixed.
  • A fixed corpus cannot represent the freshness or changing access conditions of the live web.
  • BrowseComp-Plus scores cannot be substituted for BrowseComp scores.

Explore BrowseComp-Plus · Waterloo / CSIRO / collaborators · Methodology

FrontierMath · Tiers 1–3 / Tier 4 · v2 · Source guide

Can it solve a difficult research-level math problem? Expert-authored mathematical problems with automatically checkable answers.

Measure
Solved-problem rate under the evaluator's tools and compute budget.
Comparison rule
Keep version, difficulty tier, holdout set, and tool budget explicit.
  • Tier 4 and Tiers 1–3 answer different difficulty questions.
  • OpenAI funded the original benchmark and has access to some problems; Epoch describes the held-out subsets.

Explore FrontierMath · Epoch AI · Methodology

AstaBench · Scientific research suite · Source guide

Can it carry out the steps of scientific research? Evidence spanning literature work, code execution, data analysis, and discovery.

Measure
Task-specific research outcomes across a suite of scientific evaluations.
Comparison rule
Select a named sub-benchmark and matching tools; retain agent configuration and uncertainty.
  • Not every submitted agent supports every task family.
  • A suite average can hide a strong specialty and a missing capability.

Explore AstaBench · Allen Institute for AI · Methodology

SciCode · 2024 · main problems · Source guide

Can it translate scientific knowledge into working code? Coding problems drawn from numerical methods, simulations, and scientific calculations.

Measure
Main-problem or subproblem correctness against scientific test cases.
Comparison rule
Keep background-information setting, subproblem assistance, and model-generated versus gold earlier steps separate.
  • An assisted subproblem score is not an end-to-end scientific workflow score.
  • This adds more value inside the science category than as another homepage coding total.

Explore SciCode · SciCode authors · Methodology

CritPt · Research-level physics · Source guide

Can it reason through an unfamiliar physics research problem? Challenging physics tasks that test reasoning beyond standard academic exams.

Measure
Accuracy on the evaluator's research-level physics problems.
Comparison rule
Keep the evaluation revision, tools, reasoning budget, and grading protocol fixed.
  • A specialist physics score should not stand in for general scientific usefulness.
  • Difficulty and domain coverage differ from HLE and Terminal-Bench-Science.

Explore CritPt · Artificial Analysis / CritPt authors · Methodology

LiveBench · 2026-06-25 · Source guide

How does it handle fresh, objectively scored tasks? A periodically refreshed collection covering reasoning, language, data, instruction following, and coding.

Measure
Objective task scores and category results on a named release.
Comparison rule
Compare only the same release and task categories; refreshes change the exam.
  • Its broad score overlaps with other general-purpose indices.
  • Freshness reduces contamination risk; it does not establish its absence.

Explore LiveBench · LiveBench team · Methodology

LiveCodeBench · Release v6 · dated windows · Source guide

Can it solve a new programming problem? Competitive-programming problems published over time, with executable tests.

Measure
Code-generation correctness on a specified problem-date window and scenario.
Comparison rule
Match release, start/end dates, scenario, sampling count, and execution budget.
  • Algorithmic problem solving does not establish repository-editing skill.
  • Different date windows are different cohorts, even if both are called v6.

Explore LiveCodeBench · LiveCodeBench authors · Methodology

Humanity’s Last Exam · Classic · 2025-04-03 final set · Source guide

Can it answer difficult questions across expert fields? An academic breadth check spanning mathematics, science, and the humanities, with text and image questions.

Measure
Answer accuracy on the finalized 2,500-question classic set; calibration is a separate measure.
Comparison rule
Pin the dataset revision, full versus text-only set, tools, reasoning effort, and answer judge. HLE-Rolling is a different exam.
  • Closed-ended expert questions do not measure open-ended discovery or professional work.
  • Tool-assisted and no-tool scores cannot be pooled.
  • The dataset is gated; this guide links to it without redistributing its questions.

Explore Humanity’s Last Exam · Center for AI Safety / Scale AI · Methodology

GPQA Diamond · Diamond · 198 questions · Source guide

Can it reason through an expert science question? A compact multiple-choice test in biology, chemistry, and physics, selected through expert and non-expert review.

Measure
Multiple-choice answer accuracy on the 198-question Diamond subset; random choice has a 25% baseline.
Comparison rule
Keep Diamond separate from Main and Extended. Match prompts, answer shuffling, tools, reasoning budget, and sampling policy.
  • A small fixed exam does not establish scientific discovery or laboratory competence.
  • Near-ceiling scores and small gaps require uncertainty, not a confident rank order.
  • Best-of-many success is not single-attempt accuracy.

Explore GPQA Diamond · GPQA authors · Methodology

SWE-bench Verified · Verified · 500 tasks · Source guide

Can it repair an issue in an existing repository? Real repository issues with executable tests, useful for understanding a coding agent's patching ability.

Measure
Percentage of the 500 human-filtered task instances resolved.
Comparison rule
Use one task revision, agent, budget, and attempt policy. The Bash Only view controls the agent; the full board mixes systems.
  • mini-SWE-agent 1.x and 2.x change action handling and sampling and are not automatically comparable.
  • Public task exposure and test quality limit conclusions about new, unseen work.
  • A repair score does not measure an entire software development workflow.

Explore SWE-bench Verified · SWE-bench team · Methodology

SWE-bench Pro · Public · 731 tasks · Source guide

Can it make a larger change in a complex codebase? Longer software-engineering tasks across public application and developer-tool repositories.

Measure
Resolve rate: the percentage of public tasks whose patches pass the required new and regression tests.
Comparison rule
Keep Public, Private, and Held-out sets separate. Match dataset revision, agent harness, turn limit, cost cap, and attempts.
  • The owner board mixes harnesses and capped versus uncapped runs; its rows are not one controlled cohort.
  • Public repository licensing is not evidence that models have never seen the code.
  • Task-quality disputes make this supporting evidence, not a universal replacement for Verified or Terminal-Bench.

Explore SWE-bench Pro · Scale AI · Methodology

CursorBench · 3.2 · Source guide

Which model handles the kind of work done inside Cursor? Ambiguous multi-file tasks drawn from Cursor usage, including instruction following and advanced tool use.

Measure
Cursor's task-correctness score in percent, alongside average API-priced cost per task, tokens, and steps.
Comparison rule
Compare only version 3.2 with its Cursor agent setup and exact model effort. Preserve the pricing revision for cost comparisons.
  • This is a vendor-owned internal evaluation, not an independent cross-product ranking.
  • Private tasks and agentic grading limit external reproduction.
  • The task set changed from 3.1; small score gaps may reflect evaluation variance.

Explore CursorBench · Cursor · vendor-reported · Methodology

GDPval · 2025 · original evaluation · Source guide

Can it deliver work an experienced professional would accept? Occupational tasks with reference files and finished deliverables, including documents, presentations, and spreadsheets.

Measure
Quality of completed work compared with expert deliverables across 44 occupations; the public gold set contains 220 tasks.
Comparison rule
Name the full or gold task set, grading protocol, scaffolding, and whether ties count toward the reported win rate.
  • The original evaluation and Artificial Analysis's GDPval-AA use different evaluation protocols.
  • Producing one deliverable is not the same as doing an entire job or handling its organizational context.
  • Developer-reported results need to retain their exact human or model-judge protocol.

Explore GDPval · OpenAI · Methodology

GDPval-AA · v2 · Source guide

Which tool-using model produces the strongest professional deliverable? Artificial Analysis evaluates GDPval work products in its Stirrup agent environment, then compares the outputs head to head.

Measure
Pairwise Elo rating anchored to a human-expert baseline of 1,000, with source-reported uncertainty.
Comparison rule
Keep v2, the Stirrup environment, reasoning effort, judge panel, and rating pool together. Elo is not percent correct.
  • The v2 environment, turn limit, and panel of judges differ from v1.
  • Model-judged preferences are not a direct measurement of business value or worker replacement.
  • Per-task cost must not be mixed with full-evaluation spend.

Explore GDPval-AA · Artificial Analysis · Methodology

OSWorld 2.0 · osworld-v2-2026.08.08 · Source guide

Can an agent finish a workflow across desktop and web apps? Long computer-use tasks with verifiable outcomes, not just recognizing a button in a screenshot.

Measure
Task completion and partial reward across the pinned 108-workflow release, reported separately.
Comparison rule
Pin code, tasks, assets, website, provider image, step budget, and input/action interface before comparing systems.
  • OSWorld-Verified and OSWorld 2.0 are different task cohorts.
  • Partial progress is not a completed workflow.
  • Gated environment assets and long runs affect reproducibility; the release manifest identifies the required versions.

Explore OSWorld 2.0 · OSWorld / XLang Lab · Methodology

τ³-bench · 3 · v1.0.1 grading · Source guide

Can a service agent solve the issue while following policy? Simulated customer-service conversations combine tool actions, user coordination, and domain rules; newer tracks add knowledge retrieval and voice.

Measure
Task success and repeated-trial reliability within a named domain and communication mode.
Comparison rule
Match domain, task split, user simulator, trials, and text or voice mode. pass^k consistency is not pass@k best-of-k success.
  • The repository retains the tau2-bench name while the current suite is τ³-bench.
  • Banking-knowledge results before v1.0.1 are not comparable with the corrected grading.
  • Simulated service interactions do not cover every live customer or organizational policy.

Explore τ³-bench · Sierra Research · Methodology

WISE Verified · Verified · Qwen3.5-35B-A3B · 29 charted results

Can it draw what a prompt implies? Image generation that needs world knowledge: culture, time, space, biology, physics, and chemistry.

Measure
Weighted knowledge-consistency score, 0–1; higher is better.
Comparison rule
Only the Verified prompts and Qwen3.5-35B-A3B judge. Keep agent and chain-of-thought configurations named.
Evaluation cohort
Same 1,000 Verified prompts and Qwen3.5-35B-A3B judge. Chain-of-thought and agent systems remain separate configurations. This measures world-knowledge consistency, not aesthetic preference.
Score
Knowledge consistency (score); higher is better.
Source retrieved
Source revision
sha256:1b06a1e2697244ff1f29417662bc34b5708587c6bf335e0de65c80d984972263
  • Not an aesthetic preference or image-editing test.
  • The Verified prompts and judge changed in 2026; legacy WISE scores are not comparable.
  • Published research cohort, not a complete inventory of today's image models.

Explore WISE Verified · WISE benchmark team · Methodology · Download this dataset

GEditBench · 2 · 16 charted results

Can it make an edit without breaking the rest? Instruction following, visual quality, and preservation of the original image across 23 editing tasks.

Measure
Overall pairwise Elo; higher is better. Publisher bootstrap intervals included.
Comparison rule
GPT-4o judges instruction and quality; PVC-Judge scores consistency. Compare within this fixed evaluation pool.
Evaluation cohort
GPT-4o judges instruction following and quality; PVC-Judge scores preservation. Elo and bootstrap intervals belong only to this 16-model research pool. Evaluated sample counts vary by model.
Score
Editing overall (Elo); higher is better.
Source retrieved
Source revision
sha256:87a97fa869d8d879e2f834520a109632e41c7f8e43724aabae6bb13882f838cc
  • Automated, human-aligned judging is not direct human voting.
  • This is the 2026 paper cohort; the dated API names are retained.
  • Safety refusals and failures leave some models with fewer evaluated samples.
  • Elo is pool-relative, not a percentage; overlapping intervals do not establish a winner.

Explore GEditBench · GEditBench v2 team · Methodology · Download this dataset

VideoPhy · 2 · human evaluation · 7 charted results

Does the action obey basic physics? Human reviewers check whether a generated video both follows the prompt and respects physical commonsense.

Measure
Joint semantic and physical adherence, %; higher is better.
Comparison rule
Human-evaluated All subset only. Hard, physical-activity, and object-interaction scores remain separate details.
Evaluation cohort
Percentage satisfying both semantic adherence and physical commonsense in the All subset. These 2025 human evaluations are a historical diagnostic, not a current video-model buying guide.
Score
Prompt + physical adherence (%); higher is better.
Source retrieved
Source revision
sha256:bad716a62a9247323932ca8ea305feb4a2f373ebe01233beece8e3dbea4fbd34
  • A 2025 research cohort, not a current ranking of video generators.
  • Physical plausibility is different from cinematic appeal.
  • Automatic VideoPhy2-eval scores must not be mixed with these human judgments.

Explore VideoPhy · VideoPhy2 team · Methodology · Download this dataset

OmniDocBench · 1.6_full · 35 charted results

Can it turn a difficult document into usable text? Document extraction across text, formulas, tables, and reading order, comparing specialist pipelines with general vision models.

Measure
Publisher overall score, 0–100; higher is better. Component error rates retain their native direction.
Comparison rule
Use the exact v1.6_full result table; do not pool v1.0, v1.5, or other document datasets.
Evaluation cohort
The publisher's v1.6_full table is retained even though the repository now advertises v1.7. Overall combines text, formula, and table performance. Component error metrics are lower-is-better; CDM and TEDS are higher-is-better.
Score
Document extraction overall (/ 100); higher is better.
Source retrieved
Source revision
sha256:a799f3dabee0d8e2be9f990fc325c3fefadb8d7f8d0cf56148406547d495a3ad
  • The repository advertises v1.7 but still labels this model table v1.6_full; its table label is preserved.
  • A parsing system and a general vision model are different deployment choices.
  • No matched latency or cost measurements are provided in this extract.

Explore OmniDocBench · OpenDataLab · Methodology · Download this dataset

WorldScore · Static · 2025 author cohort · 19 charted results

Can a generated world stay coherent as the camera moves? A common scene-generation protocol compares camera control, scene consistency, and quality across video, 3D, and 4D systems.

Measure
WorldScore-Static, 0–100; higher is better.
Comparison rule
Only systems sampled and evaluated by the WorldScore authors on March 30, 2025. Input and system types stay visible.
Evaluation cohort
All 19 systems were sampled and evaluated by the WorldScore authors on March 30, 2025. Static scene quality and control are compared in one protocol; this does not rank today's interactive world simulators. Newer self-evaluated submissions are excluded.
Score
WorldScore-Static (/ 100); higher is better.
Source retrieved
Source observation date
Source revision
sha256:a1c51fa9e6f1d4970f58293f4f648dda890a68d987f48b21ab8e61e0542adff3
  • Historical research comparison; newer model-team submissions are not in this chart.
  • Static and Dynamic are distinct scores; this chart does not measure interactive control latency or physical simulation accuracy.
  • A strong generated video is not evidence of an action-conditioned world simulator.

Explore WorldScore · WorldScore team · Methodology · Download this dataset

GenEval2 · 2025-12 · Source guide

Are the objects, attributes, and relationships correct? Checks the detailed compositional content of generated images, beyond whether an image looks good.

Measure
Soft-TIFA compositional alignment; higher is better.
Comparison rule
Keep arithmetic and geometric aggregation separate; GenEval2 is not the original GenEval score.
  • Judge and prompt-processing settings affect results.
  • Use the same benchmark release and scoring configuration when comparing results.

Explore GenEval2 · GenEval2 authors / Meta · Methodology

DPG-Bench · 2024-03 · Source guide

Does a detailed image prompt survive intact? Breaks dense text-to-image prompts into questions about the requested content.

Measure
Dense-prompt alignment score; higher is better.
Comparison rule
Require identical prompts, question dependencies, VQA evaluator, and rewriting policy.
  • Prompt rewriting can change what is being tested.
  • A legacy compositional test, not evidence of current aesthetic preference or editing quality.

Explore DPG-Bench · ELLA / DPG-Bench authors · Methodology

T2I-CompBench++ · ++ · 2024 suite · Source guide

Can it bind the right attribute to the right object? Separates color, shape, texture, relationships, counting, and complex scene composition.

Measure
Dimension-specific composition scores; higher is better.
Comparison rule
Use the same ++ task split and evaluator per dimension; avoid a made-up average of unlike evaluators.
  • The original suite and ++ extension have different task coverage.
  • Detector and visual-judge errors can look like generation failures.

Explore T2I-CompBench++ · T2I-CompBench authors

VBench · 2.0 · 2025 · Source guide

Where does a video generator break down? A diagnostic view of human fidelity, creativity, controllability, physics, and commonsense across 18 dimensions.

Measure
Dimension and five-aspect aggregate scores; higher is better.
Comparison rule
Keep VBench 2.0 separate from VBench 1.0, VBench++, and image-to-video tracks.
  • The overall score gives equal weight to five aspects, not to every dimension.
  • Submission settings and evaluator revisions affect comparability.
  • Automatic diagnostics do not replace viewer preference.

Explore VBench · VBench authors · Methodology

WorldModelBench · 2025 · Source guide

Does a predicted world follow the instruction and physical rules? Image-conditioned world generation judged across everyday, simulated, and embodied environments.

Measure
Instruction, commonsense, and physical-adherence scores; higher is better.
Comparison rule
Hold the input images, environment domains, and world-model evaluator fixed.
  • A small research suite is not an interactive-agent deployment test.
  • The judge and evaluated model cohort matter; a newer demo is not a benchmark result.

Explore WorldModelBench · WorldModelBench team · Methodology

Open ASR Leaderboard · Public evaluation tracks · Source guide

Which system transcribes speech accurately and quickly? Speech recognition across datasets and languages, with accuracy and inference speed reported separately.

Measure
Word error rate: lower is better. Real-time factor speedup: higher is better.
Comparison rule
Match language, dataset, chunking, decoding settings, and hardware before comparing speed.
  • Short-form, long-form, and multilingual tracks are different cohorts.
  • GPU throughput is not end-to-end API latency.
  • Hardware, chunking, and decoding settings must match for a fair efficiency comparison.

Explore Open ASR Leaderboard · Hugging Face audio team · Methodology

SEED-TTS-Eval · 2024 · Source guide

Can it say the right words in the requested voice? Zero-shot speech synthesis evaluated for intelligibility and similarity to a reference speaker in English and Mandarin.

Measure
Word error rate: lower is better. Speaker similarity: higher is better.
Comparison rule
Keep language, ASR scorer, speaker encoder, and voice-conditioning protocol fixed.
  • Speaker similarity and correct words do not measure expressive naturalness.
  • Voice-cloning evaluation is distinct from choosing a ready-made production voice.

Explore SEED-TTS-Eval · ByteDance Speech

Voice Arena · TTS v1 · Source guide

Which synthetic voice do listeners prefer? Pairwise listener preference for speech synthesis, separated by language.

Measure
Language-specific preference rating; higher is better.
Comparison rule
Compare within the same language and voice setup; retain uncertainty and vote counts.
  • Different language pools are not on one universal scale.
  • Listener preference does not establish word-perfect transcription, speaker cloning, or conversational latency.

Explore Voice Arena · Voice Arena · Methodology

MMAU-Pro · 2025-08 · Source guide

Can it reason about what it hears? Audio understanding across speech, environmental sounds, and music, including long and multiple recordings.

Measure
Task-specific audio reasoning accuracy; higher is better.
Comparison rule
Keep multiple-choice, open-ended, and instruction-following evaluation modes distinct.
  • Not a speech-generation or music-generation preference test.
  • Transcription-only systems do not receive the same sensory input as audio-native models.

Explore MMAU-Pro · MMAU-Pro authors · Methodology

MMMU-Pro · 2024 suite · Source guide

Can it reason from a diagram, not just read its text? Expert-level multimodal problems designed to reduce text-only shortcuts.

Measure
Accuracy, %; higher is better.
Comparison rule
Keep Standard 10-option and Vision settings named; original MMMU and MMMU-Pro are not interchangeable.
  • Prompting, resolution, tools, and reasoning effort affect the result.
  • Academic question answering is not document parsing or visual design quality.

Explore MMMU-Pro · MMMU authors · Methodology

Video-MME · 2024 / CVPR 2025 · Source guide

Can it understand a long video? Video question answering at short, medium, and long durations.

Measure
Question-answer accuracy, %; higher is better.
Comparison rule
Keep with-subtitle and without-subtitle tracks separate; record frame sampling and audio access.
  • Extra frames, subtitles, and audio change the information supplied to the model.
  • Understanding video is different from generating it.

Explore Video-MME · Video-MME authors · Methodology

Image Arena · Live · text-to-image · Source guide

Which generated image do people prefer? Blind human preference provides an aesthetic and overall-utility view alongside diagnostic image benchmarks.

Measure
Pairwise Elo with confidence intervals; higher is better.
Comparison rule
Text-to-image and editing have separate opponent pools; retain model settings, votes, and rating uncertainty.
  • Preference is not a factuality or exact-composition guarantee.
  • Use the publisher’s current pool, vote counts, and intervals when comparing its scores.

Explore Image Arena · Artificial Analysis

Video Arena · Live · generation tracks · Source guide

Which video looks best to viewers? Human preference for generated video, with separate input and audio tracks.

Measure
Pairwise Elo with confidence intervals; higher is better.
Comparison rule
Keep text-to-video, image-to-video, and with-audio pools separate; match resolution, duration, and frame rate for cost.
  • A silent-video score does not describe audio quality.
  • Use the publisher’s matching input, duration, resolution, and audio track when comparing results.

Explore Video Arena · Artificial Analysis · Methodology

Open ASR · meeting transcription · AMI-Cleaned · English test · 2026-09-04 · 10 charted results

Which open-weight speech models make fewer errors in English meetings? Ten selected configurations on the same cleaned meeting-transcription test. These are publisher-reported results with different inference pipelines.

Measure
Word error rate (WER), %; lower is better. Counts substituted, missing, and extra words.
Comparison rule
Compare only the AMI-Cleaned English test in this September 4, 2026 snapshot. Keep the model version and publisher scoring protocol fixed.
Evaluation cohort
Same AMI-Cleaned English test and publisher scoring protocol. Inference pipelines differ; exact per-run configurations are not published. Lower word-error rate is better. This is not a speed, speaker-attribution, or overall audio-quality ranking.
Score
Word error rate (% WER); lower is better.
Source retrieved
Source observation date
Source revision
ba5712d5ace8f785fa0daae1aecea8561ecd87c9
  • A selected open-weight comparison, not the full leaderboard or a top-ten list.
  • Inference pipelines differ. The published rows do not identify exact run dates, decoding settings, checkpoint revisions, or execution records.
  • Short speech clips do not test whole-meeting speaker attribution, punctuation quality, other languages, or live response time.
  • No uncertainty is published; small differences do not establish a reliable winner. WER can exceed 100% when extra words are inserted.

Explore Open ASR · meeting transcription · Hugging Face Open ASR Leaderboard · Methodology · Download this dataset

Terminal-Bench 4 coding standard

Terminal-Bench 4.0.0 is the site’s standard agentic terminal-engineering benchmark. Explore it in the benchmark library. The checked snapshot contains 13 configurations from the official Harbor Framework submissions at commit c7878bd, committed on , with 66 tasks and 5 trials per task. aicharts retrieved this owner snapshot on .

This standalone Terminal-Bench 4 owner cohort remains separate from the Terminal-Bench 4 component reported inside the Artificial Analysis Coding Agent Index. Every standalone TB4 row retains its model, agent, agent version, effort, accuracy, 95% confidence interval, trials, cost, tokens, duration, and pinned source files.

Download Terminal-Bench 4 JSON

Terminal-Bench-Science 0.1

The checked scientific-workflow snapshot published by Terminal-Bench-Science and Harbor Framework contains 17 owner-published system configurations across 70 tasks and 3 trials per task. It is pinned to the exact v0.1.0 release commit f81afac. The release has the persistent citation https://doi.org/10.5281/zenodo.22110254, and the owner leaderboard was updated on . aicharts retrieved it on .

Every row keeps the named model, harness, reasoning effort, resolution rate, binomial standard error, trial count, evaluation cost, token use, and owner-published source link. Terminal-Bench-Science remains separate from general terminal engineering and does not feed a composite score. Owner-published aggregate and per-domain cost fields are retained independently and are not forced to reconcile.

Download Terminal-Bench-Science 0.1 JSON

Current Intelligence efficiency · v4.3.2

The homepage’s Pareto chart uses Artificial Analysis Intelligence Index v4.3.2. Its output-token and cost views compare the identical 97-configuration positive-cost cohort from 101 complete score-and-output records. Output tokens include answer and reasoning; task cost also includes input and cache traffic.

This 10-evaluation version weights agents 30%, coding 20%, scientific reasoning 20%, and general capability 30%. The source is checked every four hours. Version, evaluation roster, source identities, native measures, and retention must pass validation before an update is published. Retrieved .

Download current Intelligence v4.3.2 JSON · Full v4.3.2 comparison rules and limitations. Keep these results separate from the frozen v4.1.1 snapshot below; its evaluation roster and weights differ.

Historical Intelligence v4.1.1 · frozen snapshot

This retained historical dataset pairs the owner-published Artificial Analysis Intelligence Index v4.1.1 score with weighted output tokens and cost per Intelligence Index task. It is not the current homepage dataset. The frozen snapshot retains 135 measured score-and-output records from 643 source records. Its historical matched-resource comparison uses the same 127-record cohort with positive comparable cost.

The Index combines 9 evaluations. Its four owner-defined category weights and constituent evaluation weights are:

  • Agents · 34%: GDPval-AA v2 · 20%; τ³-Banking · 14%.
  • Coding · 24%: Terminal-Bench v2.1 · 16%; SciCode · 8%.
  • Scientific Reasoning · 24%: Humanity's Last Exam · 12%; GPQA Diamond · 6%; CritPt · 6%.
  • General · 18%: AA-LCR · 6%; AA-Omniscience · 12%.

Output tokens here mean answer plus reasoning tokens only, weighted by each evaluation's Index weight and divided by its task count. They are not the coding-agent chart's total tokens, which also include input traffic. Cost is the owner's weighted per-task sum of available input, cache, reasoning, and answer/output components. A source row with a complete cost breakdown but a reported zero total is stored as unavailable, never converted into a free-model value. Rows with incomplete cost are excluded. Complete zero-total rows remain in the JSON but are omitted from the historical matched-resource cohort.

The comparable cohort follows the checked rule: current, non-estimated model configurations with a finite 0–100 Intelligence Index, finite positive per-task output-token total, and a complete finite nonnegative per-task cost breakdown; source cost totals at or below zero normalize to null. In a Pareto frontier for that stored cohort, a frontier point is not dominated by another record with an equal-or-higher Intelligence score and equal-or-lower output-token or positive-cost value. Artificial Analysis publishes the measurements; the frontier classification is aicharts analysis.

This v4.1.1 snapshot is frozen and is no longer refreshed by automation. The current v4.3.2 dataset above has a separate versioned source contract and download. aicharts retrieved this historical snapshot on .

Read the owner's Index methodology, model leaderboard, and terms of use. Citation: Artificial Analysis (2025). LLM benchmarks dataset. https://artificialanalysis.ai.

Download historical v4.1.1 JSON

Artificial Analysis coding-agent source and refresh

The source is the public Coding Agent Index v1.5 Artificial Analysis coding-agents comparison, composed of DeepSWE v1.1, Terminal-Bench 4, and SWE-Atlas-QnA. aicharts retrieved this snapshot on . The site checks for a new source snapshot daily. The displayed retrieval time changes only when a validated snapshot is stored.

The most recent retained model, variant, or material benchmark change was detected on . This meaningful-update time is separate from the daily retrieval check.

The checked dataset contains 20 model-agent configurations across 19 models, 8 agent harnesses, and 10 model providers. The chart and the JSON download use this same checked snapshot.

Benchmark definitions

Each benchmark is shown on the 0–100 scale stored in the snapshot. The metrics evaluate different tasks and should be interpreted separately.

AA Index

Overall performance across code changes, terminal work, and repository understanding.

DeepSWE v1.1

Long-horizon software engineering tasks scored with automated code verification.

Terminal-Bench 4

Agentic terminal-use tasks scored with automated test-suite verification.

SWE-Atlas-QnA

Repository-understanding questions scored with a strict resolve verifier.

Current leaders

These are the highest available scores in the retrieved snapshot, one row per benchmark. They are observations of the named model, agent harness, and effort setting rather than general model ranks.

Highest score by benchmark in the current snapshot
BenchmarkModelAgentProviderSettingScore
AA IndexOpus 5.5Claude CodeAnthropicmax66.0
DeepSWE v1.1Muse Spark 1.3Muse CodeMetaxhigh73.2
Terminal-Bench 4Opus 5.5Claude CodeAnthropicmax63.1
SWE-Atlas-QnAOpus 5.5Claude CodeAnthropicmax66.4

For AA Index versus mean API cost, including the cost/performance frontier, see highest AA Index and lowest cost pick different agents. For whether classified open-weight rows sit with those leaders, see open models closed SemiAnalysis composites, not this table. For how a cheaper model changed one daily news page, see GPT-5.6 Luna made one daily news page cost about $0.10. For what a 30% Terminal-Bench-Science result measures, and how cost and token use change the comparison, see What Terminal-Bench-Science’s 30% result measures. For why a public-suite high score still needs a holdout, see why a coding-agent high score still needs a holdout.

All configurations

Every model-agent configuration in the retrieved snapshot, with AA Index, component scores, and mean API cost per task. Missing values are stored as empty in the source and shown as a dash.

All 20 model-agent configurations in the Artificial Analysis snapshot retrieved Sep 25, 2026, 2:39 PM UTC
ModelAgentProviderSettingAA IndexDeepSWE v1.1Terminal-Bench 4SWE-Atlas-QnACost
Opus 5.5Claude CodeAnthropicmax66.068.463.166.4$13.04
Fable 5.1 (with fallback)Claude CodeAnthropicmax62.264.357.664.8$12.39
Claude Fable 5.1 XHigh + SWE-2 MediumDevin Fusion CLICognitiondefault61.763.156.165.9$7.90
GPT-6 AstraCodexOpenAImax61.667.655.661.8$7.47
Opus 5Claude CodeAnthropicmax59.762.554.562.1$10.79
GPT-6 Astra XHigh + SWE-2 MediumDevin Fusion CLICognitiondefault58.967.350.059.4$4.54
GPT-6 SolCodexOpenAImax56.769.043.457.5$2.99
Grok 4.7Grok BuildxAIxhigh56.372.633.362.9$8.82
GPT-5.6 SolCodexOpenAImax54.672.337.454.0$6.35
Muse Spark 1.3Muse CodeMetamax54.371.731.859.4$3.98
GLM-5.3OpencodeZ.aidefault53.661.439.959.4$4.24
Kimi K3Kimi Code CLIMoonshot AIdefault51.968.421.266.1$5.05
Muse Spark 1.3Muse CodeMetaxhigh48.373.217.254.6$3.47
Grok 4.6Grok BuildxAIxhigh47.064.917.758.3$3.57
Qwen3.8 MaxClaude CodeAlibaba Clouddefault43.351.016.762.1$3.48
GPT-5.6 LunaCodexOpenAImax43.266.414.648.7$0.438
DeepSeek V4 Pro 0813CodexDeepSeekmax43.157.210.161.8$0.238
Gemini 3.8 FlashAntigravity SDKGooglehigh41.965.814.645.2$2.47
GPT-6 LunaCodexOpenAImax41.163.715.244.4$0.176
DeepSeek V4 Flash 0731CodexDeepSeekmax38.754.310.651.3$0.085

Normalization method

The refresh job reads the source page's public data payload, validates every source row, and maps it into a versioned owned schema. Benchmark reward proportions are represented as 0–100 scores. Mean task cost stays in US dollars, mean active wall time stays in seconds, and mean total token use stays as a token count.

Provider identifiers, model effort settings, stable series keys, and sort order are normalized for the chart. The refresh is rejected when duplicate records, major row loss, stable-key loss, or substantial metric-coverage regressions are detected. aicharts does not recalculate the source benchmark outcomes.

Limitations

  • Artificial Analysis defines and operates the upstream evaluations. aicharts is an independent visualization and is not affiliated with Artificial Analysis or the listed providers.
  • The Intelligence efficiency view is an owner-defined, primarily English-language aggregate. Its category weights emphasize agentic tasks, and it does not establish performance for every use case.
  • Scores depend on the named model, agent harness, effort setting, task set, and evaluation version. They do not establish results for every software repository or production workflow.
  • Cost, duration, and token values are task-level means from the source evaluation. They are not price or latency guarantees.
  • The current v4.3.2 Intelligence snapshot is checked every four hours and the coding-agent snapshot daily; neither is a real-time mirror. The historical v4.1.1 snapshot is frozen. Use the relevant version and retrieval timestamp when citing a value.
  1. Download the current Intelligence v4.3.2 JSON snapshotModel-level Intelligence score, output-only tokens, comparable cost, source method, version, and retrieval time.
  2. Historical Intelligence v4.1.1 JSONFrozen earlier cohort; not refreshed or comparable with the current index scale.
  3. Artificial Analysis model leaderboardThe upstream model comparison; its public structured data is a source-shape cross-check rather than the full snapshot source.
  4. Current Intelligence refresh and normalization source codeThe public parser, source cross-checks, normalization rules, and replacement guards.
  5. Download the coding-agent JSON snapshotVersioned records, provenance, retrieval time, and bounded update history used by the production chart.
  6. Artificial Analysis coding-agents sourceThe upstream comparison from which the checked snapshot is derived.
  7. Refresh and normalization source codeThe public parser, normalization rules, validation guards, and update-detection logic.