Benchmark atlas: charts and source guides
The atlas covers 62 benchmarks and 29 charted evaluation cohorts. Each chart uses one source, score definition, and version. A source guide explains an evaluation whose results are not yet charted here. Historical research cohorts remain labeled; retrieving a paper today does not mean its models were evaluated today.
Download the catalog JSON to discover the available benchmark IDs, comparison rules, source dates, and individual dataset URLs. Each dataset download includes the same observations, units, costs, uncertainty labels, and configurations used by its chart. This is a checked snapshot distribution, not a live model API.
Cite the benchmark owner and source version when quoting measurements. aicharts publishes these chart projections; third-party measurements retain their source terms. The software license does not grant a new license to third-party data.
Terminal-Bench 4 · 4.0.0 · 13 charted results
Which agent can finish difficult terminal work? Software, systems, CAD, scientific computing, and formal proof tasks executed in a terminal.
- Measure
- Task success over 66 tasks and five trials per configuration, with source-reported 95% confidence intervals.
- Comparison rule
- aicharts’ primary terminal benchmark. Compare exact version 4.0.0 with the named model, agent version, and effort.
- Evaluation cohort
- 66 tasks × 5 trials per configuration. Model, agent, and effort all affect results. The owner reports 95% confidence intervals; its interval method and price basis are unspecified.
- Score
- Task success (%); higher is better.
- Source retrieved
- Source revision
c7878bd7cbd05d1dd21586b83eb3a1929a412312- Cost basis
- USD for the full 330-trial evaluation
- An agent and model are evaluated together; this is not a model-only ranking.
- Different harnesses and effort settings remain separate configurations.
- Cost covers the full 330-trial evaluation. The source does not specify its pricing or confidence-interval method.
Explore Terminal-Bench 4 · Harbor Framework · Methodology · Download this dataset
Artificial Analysis Intelligence Index · 4.3.2 · 101 charted results
How much capability do I get for the cost? Broad model capability alongside the cost and output tokens used to achieve it, within the same evaluation cohort.
- Measure
- The publisher’s v4.3.2 index across ten evaluations: agents 30%, coding 20%, scientific reasoning 20%, and general capability 30%.
- Comparison rule
- Compare configurations only within v4.3.2. Its evaluation roster and weights differ from v4.1.1; the historical scores are not on a continuous scale with these scores.
- Evaluation cohort
- Ten-evaluation v4.3.2 cohort. Compare the same reasoning configuration and native per-task measurements. Historical v4.1.1 scores remain separate; a later index revision requires a new versioned cohort.
- Score
- Intelligence Index (index points); higher is better.
- Source retrieved
- Cost basis
- USD per Intelligence Index task
- An index reflects the publisher’s task mix and weights, not every use case. Reasoning effort changes the evaluated configuration.
- Both resources are publisher-reported per-task measures. Output tokens include answer and reasoning; cost also includes input and cache traffic.
- Only current, non-estimated configurations with complete native measurements are retained. Missing positive cost is not free. No per-configuration uncertainty is supplied.
- The Terminal-Bench component uses Artificial Analysis’s own evaluation harness and must not be merged with the Harbor submission leaderboard.
Explore Artificial Analysis Intelligence Index · Artificial Analysis · Methodology · Download this dataset
Image generation · Arena · Overall · Bradley–Terry · 80 charted results
Which image generators lead Arena’s preference ratings? Compare published preference ratings for images generated from text, with uncertainty and vote counts.
- Measure
- Published Bradley–Terry preference rating; higher is better. Source confidence intervals and vote counts included.
- Comparison rule
- Text-to-image, overall category only. One complete, dated publisher cohort; no blending with other Arena tracks or diagnostic benchmarks.
- Evaluation cohort
- Text-to-image overall ratings, published 2026-09-24. Compare within this dated pool; retain settings and confidence intervals. Preliminary and AutoEval flags are not provided. Arena (lmarena-ai), Arena leaderboard dataset, CC-BY-4.0 (https://creativecommons.org/licenses/by/4.0/). Selected overall media cohorts and projected fields for AI Charts; published ratings, intervals, and vote counts unchanged.
- Score
- Arena preference rating (Arena points); higher is better.
- Source retrieved
- Source observation date
- Source revision
6931a5ef0714c8d4edf09d79998fef6674ed4e09
- Ratings depend on the opponent pool and prompt mix. Compare only within this track and publication date; ratings are not percentages.
- Overlapping confidence intervals do not establish a clear winner. Vote count is per model, not a count of unique people or prompts.
- The licensed export omits preliminary and AutoEval status flags. A row’s evidence status cannot be inferred from its vote count.
- This extract has no matched price or latency data. Keep audio, resolution, effort, web-search, and agent settings in the model labels.
- Arena supports AutoEval proxy votes for image generation; the export does not identify affected rows. Preference does not guarantee correct composition or faithful editing.
Explore Image generation · Arena · Arena · Methodology · Download this dataset
Image editing · Arena · Overall · Bradley–Terry · 56 charted results
Which image editors lead Arena’s preference ratings? A separate preference comparison for editing an existing image, preserving each published model setting.
- Measure
- Published Bradley–Terry preference rating; higher is better. Source confidence intervals and vote counts included.
- Comparison rule
- Image editing, overall category only. One complete, dated publisher cohort; no blending with other Arena tracks or diagnostic benchmarks.
- Evaluation cohort
- Image editing overall ratings, published 2026-09-22. Compare within this dated pool; retain settings and confidence intervals. Preliminary and AutoEval flags are not provided. Arena (lmarena-ai), Arena leaderboard dataset, CC-BY-4.0 (https://creativecommons.org/licenses/by/4.0/). Selected overall media cohorts and projected fields for AI Charts; published ratings, intervals, and vote counts unchanged.
- Score
- Arena preference rating (Arena points); higher is better.
- Source retrieved
- Source observation date
- Source revision
6931a5ef0714c8d4edf09d79998fef6674ed4e09
- Ratings depend on the opponent pool and prompt mix. Compare only within this track and publication date; ratings are not percentages.
- Overlapping confidence intervals do not establish a clear winner. Vote count is per model, not a count of unique people or prompts.
- The licensed export omits preliminary and AutoEval status flags. A row’s evidence status cannot be inferred from its vote count.
- This extract has no matched price or latency data. Keep audio, resolution, effort, web-search, and agent settings in the model labels.
- Arena supports AutoEval proxy votes for image generation; the export does not identify affected rows. Preference does not guarantee correct composition or faithful editing.
Explore Image editing · Arena · Arena · Methodology · Download this dataset
Video generation · Arena · Overall · Bradley–Terry · 48 charted results
Which text-to-video systems lead Arena’s preference ratings? Compare video generators in the same text-prompt preference pool, keeping audio, resolution, and agent labels intact.
- Measure
- Published Bradley–Terry preference rating; higher is better. Source confidence intervals and vote counts included.
- Comparison rule
- Text-to-video, overall category only. One complete, dated publisher cohort; no blending with other Arena tracks or diagnostic benchmarks.
- Evaluation cohort
- Text-to-video overall ratings, published 2026-09-22. Compare within this dated pool; retain settings and confidence intervals. Preliminary and AutoEval flags are not provided. Arena (lmarena-ai), Arena leaderboard dataset, CC-BY-4.0 (https://creativecommons.org/licenses/by/4.0/). Selected overall media cohorts and projected fields for AI Charts; published ratings, intervals, and vote counts unchanged.
- Score
- Arena preference rating (Arena points); higher is better.
- Source retrieved
- Source observation date
- Source revision
6931a5ef0714c8d4edf09d79998fef6674ed4e09
- Ratings depend on the opponent pool and prompt mix. Compare only within this track and publication date; ratings are not percentages.
- Overlapping confidence intervals do not establish a clear winner. Vote count is per model, not a count of unique people or prompts.
- The licensed export omits preliminary and AutoEval status flags. A row’s evidence status cannot be inferred from its vote count.
- This extract has no matched price or latency data. Keep audio, resolution, effort, web-search, and agent settings in the model labels.
- Video preference does not measure physical plausibility, interactive world simulation, or camera-control accuracy.
Explore Video generation · Arena · Arena · Methodology · Download this dataset
Image-to-video · Arena · Overall · Bradley–Terry · 48 charted results
Which systems turn an image into a preferred video? An image-conditioned video preference comparison, separate from generation that starts with a text prompt alone.
- Measure
- Published Bradley–Terry preference rating; higher is better. Source confidence intervals and vote counts included.
- Comparison rule
- Image-to-video, overall category only. One complete, dated publisher cohort; no blending with other Arena tracks or diagnostic benchmarks.
- Evaluation cohort
- Image-to-video overall ratings, published 2026-09-22. Compare within this dated pool; retain settings and confidence intervals. Preliminary and AutoEval flags are not provided. Arena (lmarena-ai), Arena leaderboard dataset, CC-BY-4.0 (https://creativecommons.org/licenses/by/4.0/). Selected overall media cohorts and projected fields for AI Charts; published ratings, intervals, and vote counts unchanged.
- Score
- Arena preference rating (Arena points); higher is better.
- Source retrieved
- Source observation date
- Source revision
6931a5ef0714c8d4edf09d79998fef6674ed4e09
- Ratings depend on the opponent pool and prompt mix. Compare only within this track and publication date; ratings are not percentages.
- Overlapping confidence intervals do not establish a clear winner. Vote count is per model, not a count of unique people or prompts.
- The licensed export omits preliminary and AutoEval status flags. A row’s evidence status cannot be inferred from its vote count.
- This extract has no matched price or latency data. Keep audio, resolution, effort, web-search, and agent settings in the model labels.
- Video preference does not measure physical plausibility, interactive world simulation, or camera-control accuracy.
Explore Image-to-video · Arena · Arena · Methodology · Download this dataset
Artificial Analysis Intelligence Index · 4.1.1 · 135 charted results
How much capability do I get for the cost? A versioned snapshot of broad model capability, with the output tokens and cost used to achieve it.
- Measure
- An owner-defined index across nine evaluations of agents, coding, scientific reasoning, and general knowledge.
- Comparison rule
- Compare models within the retained v4.1.1 cohort. Later index versions use different evaluations or weights and require a separate chart.
- Evaluation cohort
- Historical v4.1.1 snapshot of the nine-evaluation index. The publisher introduced v4.3 on September 7; those results are not merged here. Effort settings and missing costs remain explicit.
- Score
- Intelligence Index (index points); higher is better.
- Source retrieved
- Cost basis
- USD per Intelligence Index task
- The publisher introduced v4.3 on September 7, 2026. This historical v4.1.1 chart is not a current-version ranking.
- The index reflects the publisher’s task mix and weights, not every use case.
- Reasoning effort changes the system being evaluated. Missing cost does not mean free.
Explore Artificial Analysis Intelligence Index · Artificial Analysis · Methodology · Download this dataset
Terminal-Bench-Science · 0.1.0 · 17 charted results
Can an agent carry out scientific work? Research tasks that require scientific software, computation, and evidence across five domains.
- Measure
- Resolution rate across 70 tasks and three trials, with separate domain results.
- Comparison rule
- Compare the exact 0.1.0 cohort. Error bars show one standard error, which is narrower than a 95% confidence interval.
- Evaluation cohort
- 70 science tasks × 3 trials. Error bars show one binomial standard error, not a 95% confidence interval. Model and agent versions, price basis, and several run limits are unspecified.
- Score
- Resolution rate (%); higher is better.
- Source retrieved
- Source observation date
- Source revision
f81afac4f11048e77a15dfc8fb1dbfb897fea0ce- Cost basis
- USD for the full 210-trial evaluation
- Model and agent versions and several run limits are not specified by the source.
- Cost covers the full 210-trial evaluation; source aggregate and domain costs can differ.
- A benchmark success does not establish the scientific validity of arbitrary new research.
Explore Terminal-Bench-Science · Terminal-Bench-Science · Methodology · Download this dataset
DeepSWE v1.1 · AA evaluation · 1.1 · 20 charted results
Can an agent complete long software changes? Long-horizon software engineering, measured inside Artificial Analysis’s coding-agent evaluation.
- Measure
- Success on software engineering tasks with automated code verification.
- Comparison rule
- Keep this Artificial Analysis cohort separate from DataCurve’s direct mini-swe-agent leaderboard and other harnesses.
- Evaluation cohort
- Each point is an Artificial Analysis coding-agent configuration. Cost covers the source’s coding suite, not an isolated run of the selected benchmark. Component results stay separate from standalone cohorts published by other evaluators.
- Score
- DeepSWE v1.1 (%); higher is better.
- Source retrieved
- Cost basis
- USD per task across the AA coding suite
- Cost and time describe the AA coding suite, not the isolated DeepSWE task set.
- The checked source does not report per-configuration uncertainty.
Explore DeepSWE v1.1 · AA evaluation · Artificial Analysis · Download this dataset
Artificial Analysis Coding Agent Index · 1.5 · 20 charted results
Which coding configuration covers a range of tasks? The publisher’s combined view of software changes, terminal work, and repository understanding.
- Measure
- The source’s Coding Agent Index alongside mean evaluation cost, time, and token use.
- Comparison rule
- Read the source-defined index within this coding-suite snapshot. It is distinct from the general Intelligence Index.
- Evaluation cohort
- Each point is an Artificial Analysis coding-agent configuration. Cost covers the source’s coding suite, not an isolated run of the selected benchmark. Component results stay separate from standalone cohorts published by other evaluators.
- Score
- Coding Agent Index (index points); higher is better.
- Source retrieved
- Cost basis
- USD per task across the AA coding suite
- The combined score reflects the evaluator’s task mix.
- Different agent harnesses and effort settings affect both performance and resource use.
- Component scores remain specific to this Artificial Analysis cohort.
Explore Artificial Analysis Coding Agent Index · Artificial Analysis · Download this dataset
Terminal-Bench 4 · AA evaluation · 4.0 · 20 charted results
How did these coding configurations perform on terminal tasks? Terminal-Bench 4 results retained inside the Artificial Analysis coding-suite comparison.
- Measure
- Task success on Terminal-Bench 4 with the named agent and model.
- Comparison rule
- Keep this Artificial Analysis cohort separate from the standalone owner-published Terminal-Bench 4 cohort.
- Evaluation cohort
- Each point is an Artificial Analysis coding-agent configuration. Cost covers the source’s coding suite, not an isolated run of the selected benchmark. Component results stay separate from standalone cohorts published by other evaluators.
- Score
- Terminal-Bench 4 (%); higher is better.
- Source retrieved
- Cost basis
- USD per task across the AA coding suite
- Cost and time cover the AA coding suite, not this isolated benchmark.
- The checked source does not report per-configuration uncertainty.
Explore Terminal-Bench 4 · AA evaluation · Artificial Analysis · Download this dataset
SWE-Atlas-QnA · swe-atlas-qna · 20 charted results
Can an agent understand a software repository? Repository questions evaluated with a strict answer verifier.
- Measure
- Successful answers to repository-understanding questions in the AA evaluation.
- Comparison rule
- Compare the retained SWE-Atlas-QnA source cohort and exact agent configuration.
- Evaluation cohort
- Each point is an Artificial Analysis coding-agent configuration. Cost covers the source’s coding suite, not an isolated run of the selected benchmark. Component results stay separate from standalone cohorts published by other evaluators.
- Score
- SWE-Atlas-QnA (%); higher is better.
- Source retrieved
- Cost basis
- USD per task across the AA coding suite
- Repository understanding is narrower than successfully changing and shipping software.
- The source does not expose a finer immutable dataset revision in this snapshot.
- Cost and time cover the AA coding suite, not this isolated benchmark.
Explore SWE-Atlas-QnA · Artificial Analysis · Download this dataset
Vals Index · Vals AI · 2 · private test set · 66 charted results
Which model does the most economically valuable work? A GDP-weighted average of agentic performance across finance, coding, and legal tasks.
- Measure
- The publisher's composite of seven evaluations, weighted by each sector's share of U.S. GDP: finance 8.0, coding 5.6, and legal 1.2, over a denominator of 14.8.
- Comparison rule
- Compare within index version 2. Versions 1 and 2 use different component benchmarks, so their scores are not a series.
- Evaluation cohort
- One publisher-defined composite. Cost is the average dollar cost of one test across the weighted components, not the cost of a task you would run. Model names are the identifiers Vals publishes, not aicharts canonical names.
- Score
- Accuracy (%); higher is better.
- Source retrieved
- Source observation date
- Source revision
vals_index v2- Cost basis
- Average cost per test (USD)
- This is the publisher's composite, not an aicharts ranking. The sector weights are a deliberate simplification of how AI reaches the economy.
- The Code Migration component scores a fixed 60-task subset of the published 120-task run, not the full benchmark.
- Five of the seven components are private evaluations that no independent party can reproduce.
- Model names are the identifiers Vals publishes, not aicharts canonical names.
Explore Vals Index · Vals AI · Vals AI · Methodology · Download this dataset
Finance Agent (v2) · Vals AI · 2 · private test set · 70 charted results
Can an agent do a financial analyst's work? Multi-step analyst tasks covering modeling, comparables, earnings, and disclosure reading.
- Measure
- Severity-weighted partial credit across ten task families in the FAB v2 harness, averaged over repeated runs.
- Comparison rule
- Compare within Finance Agent v2. The v1 board used a different task set and is not a continuation of this one.
- Evaluation cohort
- Private evaluation run by Vals. Dollar values are the publisher's average cost per test under its own harness. Model names are the identifiers Vals publishes, not aicharts canonical names.
- Score
- Accuracy (%); higher is better.
- Source retrieved
- Source observation date
- Source revision
finance_agent v2- Cost basis
- Average cost per test (USD)
- The test set is private, so no independent party can reproduce these scores.
- Task-family scores in the inspector are components of the headline number, not separate rankings.
- Model names are the identifiers Vals publishes, not aicharts canonical names.
Explore Finance Agent (v2) · Vals AI · Vals AI · Methodology · Download this dataset
Legal Research Bench · Vals AI · 1 · private test set · 69 charted results
Can an agent do legal research it can cite? Case and statute research across eight areas of US law, scored against citation-backed answers.
- Measure
- Accuracy on research questions across administrative, business, civil, constitutional, criminal, family, health, and immigration law.
- Comparison rule
- Compare within Legal Research Bench version 1. Per-area scores are components, not separate leaderboards.
- Evaluation cohort
- Private evaluation run by Vals with partner law firms. Dollar values are the publisher's average cost per test. Model names are the identifiers Vals publishes, not aicharts canonical names.
- Score
- Accuracy (%); higher is better.
- Source retrieved
- Source observation date
- Source revision
legal_research v1- Cost basis
- Average cost per test (USD)
- The test set is private and was built with partner law firms, so no independent party can reproduce these scores.
- US law only. The result says nothing about research in another jurisdiction.
- A citation-backed answer that scores well is not legal advice and has not been checked by a lawyer for a specific matter.
- Model names are the identifiers Vals publishes, not aicharts canonical names.
Explore Legal Research Bench · Vals AI · Vals AI · Methodology · Download this dataset
Tax Agent Bench · Vals AI · 1 · private test set · 61 charted results
Can an agent answer a research-grade tax question? US corporate tax questions covering fact patterns, rule lookup, calculations, forms, and controversy.
- Measure
- Accuracy across six tax task families, with a stricter all-pass variant reported alongside the headline score.
- Comparison rule
- Compare within Tax Agent Bench version 1. This board covers 22 configurations, fewer than the other Vals boards.
- Evaluation cohort
- Private evaluation run by Vals. Dollar values are the publisher's average cost per test. Model names are the identifiers Vals publishes, not aicharts canonical names.
- Score
- Accuracy (%); higher is better.
- Source retrieved
- Source observation date
- Source revision
tax_agent_bench v1- Cost basis
- Average cost per test (USD)
- The test set is private, so no independent party can reproduce these scores.
- US corporate tax only, and tax rules change with the filing year.
- A score here is not tax advice and does not establish that an answer is filing-ready.
- Model names are the identifiers Vals publishes, not aicharts canonical names.
Explore Tax Agent Bench · Vals AI · Vals AI · Methodology · Download this dataset
MedCode · Vals AI · 1 · private test set · 102 charted results
Can a model assign the right medical billing code? Medical coding for the billing process, scored one-shot rather than as an agent.
- Measure
- Accuracy on medical billing code assignment in a single-response setting.
- Comparison rule
- One-shot only. These scores are not comparable with the agentic Vals boards, which let a model take many steps.
- Evaluation cohort
- Private one-shot evaluation run by Vals. The publisher does not present a comparable cost for this board. Model names are the identifiers Vals publishes, not aicharts canonical names.
- Score
- Accuracy (%); higher is better.
- Source retrieved
- Source observation date
- Source revision
medcode v1
- Vals does not present a comparable cost for this board, so no cost axis is offered.
- The board retains configurations from 2025 alongside current models, so the range spans more than one model generation.
- The test set is private, so no independent party can reproduce these scores.
- Model names are the identifiers Vals publishes, not aicharts canonical names.
Explore MedCode · Vals AI · Vals AI · Methodology · Download this dataset
ARC-AGI-2 · 2 · semi-private · 28 charted results
Can it solve an unfamiliar visual puzzle? Infer a rule from a few examples, then apply it to a new grid.
- Measure
- Percentage of novel abstract grid tasks solved in ARC Prize's semi-private evaluation.
- Comparison rule
- Compare the same dataset split and reasoning effort. This view selects seven recent model families.
- Evaluation cohort
- Semi-private set. Seven selected recent model families; cost is unpublished for these configurations.
- Score
- Tasks solved (%); higher is better.
- Source retrieved
- Source observation date
- Source revision
SHA-256 b5cb5ec6e8c7547c4f11a4e61e5ed7e0f28a26208f87bd16c83d1703777a2627
- Strong puzzle performance does not establish general intelligence.
- Selected configurations have no published cost in this source.
- Scores near the ceiling leave less room to distinguish leading systems.
Explore ARC-AGI-2 · ARC Prize · Methodology · Download this dataset
ARC-AGI-3 · Standard · 3 · semi-private · Standard · 12 charted results
Can it learn the rules by interacting? Explore an unfamiliar environment and solve it through a shared, minimal agent interface.
- Measure
- Action-efficiency score relative to the benchmark's human baseline, expressed as a percentage.
- Comparison rule
- Standard harness only. Compare the selected Astra, Sol, and Opus 5 configurations within this view.
- Evaluation cohort
- Standard harness only. Dollar values cover the full evaluation. Harness cohorts are shown separately.
- Score
- Action-efficiency score (%); higher is better.
- Source retrieved
- Source observation date
- Source revision
SHA-256 d743d7a731ba6b9f102d3624d5fe82011630b90da4200ceb34bd56cf2b221e1c- Cost basis
- Total evaluation cost (USD)
- Cost is the full evaluation spend, not the cost of one task.
- These are deterministic puzzle environments; they do not measure open-ended real-world competence.
- The Provider Adapter harness is a separate comparison.
Explore ARC-AGI-3 · Standard · ARC Prize · Methodology · Download this dataset
ARC-AGI-3 · Provider Adapter · 3 · semi-private · Provider Adapter · 6 charted results
How much does native context management help? Astra's reasoning settings with its provider's conversation and compaction features enabled.
- Measure
- ARC-AGI-3 action-efficiency score using the Provider Adapter harness.
- Comparison rule
- Compare Astra's six effort settings here. These scores are not a ranking against Standard-harness runs.
- Evaluation cohort
- Provider Adapter harness only. Dollar values cover the full evaluation. Harness cohorts are shown separately.
- Score
- Action-efficiency score (%); higher is better.
- Source retrieved
- Source observation date
- Source revision
SHA-256 d743d7a731ba6b9f102d3624d5fe82011630b90da4200ceb34bd56cf2b221e1c- Cost basis
- Total evaluation cost (USD)
- Only Astra is included in this selected adapter cohort.
- Cost is total evaluation spend.
- Near-perfect performance on this bounded test is not proof of general intelligence.
Explore ARC-AGI-3 · Provider Adapter · ARC Prize · Methodology · Download this dataset
DeepResearch Bench II · II · 132 tasks · 9 charted results
Which research agent produces a useful report? Compare information gathering, analysis, and presentation against expert-written research rubrics.
- Measure
- Owner-reported weighted rubric score across 132 tasks and 9,430 criteria; component scores appear in the inspector.
- Comparison rule
- Keep the exact research product and model generation. An o3-era research run does not represent today's OpenAI product.
- Evaluation cohort
- Nine selected systems. The table includes older products; run dates, exact model versions, costs, and confidence intervals are not consistently supplied.
- Score
- Weighted rubric score (%); higher is better.
- Source retrieved
- Source revision
SHA-256 ba0ccb0a89ff6c8f5d3c97861b072cdb8459c27fdaef2db90e725537f2869562
- The source includes historical products and does not date every run.
- A model judge evaluates reports; no intervals or comparable costs are published in this table.
- This is a selected set of nine systems, not the full source leaderboard.
Explore DeepResearch Bench II · USTC / Metastone · Methodology · Download this dataset
LongMemEval-V2 · Small · 2 · small · paper baselines · 6 charted results
Can an agent reuse what it learned before? Compare memory systems that turn previous web-agent work into useful evidence for a later question.
- Measure
- Answer accuracy with a fixed reader. Query latency is shown for each memory method.
- Comparison rule
- Compare memory methods within the same history tier and reader configuration. Small and Medium are separate sets.
- Evaluation cohort
- Small history tier only. This compares six memory methods, not six foundation models. Latency is in the inspector; dollar cost is not published.
- Score
- Answer accuracy (%); higher is better.
- Source retrieved
- Source revision
Paper baselines; evaluation code 2cc8c540bdb87fe6761629b585e727e1c4704520
- These are six paper baselines; the public submission leaderboard is still empty.
- This measures memory-system plus reader performance, not a general model ranking.
- Query latency excludes the wider application experience; no comparable dollar costs are published.
Explore LongMemEval-V2 · Small · UCLA LongMemEval team · Methodology · Download this dataset
LongMemEval-V2 · Medium · 2 · medium · paper baselines · 6 charted results
Can an agent reuse what it learned before? Compare memory systems that turn previous web-agent work into useful evidence for a later question.
- Measure
- Answer accuracy with a fixed reader. Query latency is shown for each memory method.
- Comparison rule
- Compare memory methods within the same history tier and reader configuration. Small and Medium are separate sets.
- Evaluation cohort
- Medium history tier only. This compares six memory methods, not six foundation models. Latency is in the inspector; dollar cost is not published.
- Score
- Answer accuracy (%); higher is better.
- Source retrieved
- Source revision
Paper baselines; evaluation code 2cc8c540bdb87fe6761629b585e727e1c4704520
- These are six paper baselines; the public submission leaderboard is still empty.
- This measures memory-system plus reader performance, not a general model ranking.
- Query latency excludes the wider application experience; no comparable dollar costs are published.
Explore LongMemEval-V2 · Medium · UCLA LongMemEval team · Methodology · Download this dataset
LongMemEval · 1 · S / M · Source guide
Will it remember changing facts across conversations? Tests recall, updates, time, reasoning across sessions, and knowing when an answer is absent.
- Measure
- Answer accuracy on 500 questions across long conversation histories.
- Comparison rule
- Keep S and M separate; match reader, judge, retrieval budget, and history construction.
- Retrieval recall@k is not answer accuracy.
- Vendor memory-system scores often use different readers and judges.
LoCoMo · ACL 2024 · public 10-conversation set · Source guide
Can it recall details from a long-running conversation? A widely cited test for conversational memory, event summaries, and dialogue continuity.
- Measure
- Question-answering and conversation tasks with long, multi-session histories.
- Comparison rule
- Name the public subset, included question categories, scoring method, reader, and retrieval budget.
- F1, model-judged correctness, and retrieval recall are different scores.
- The small conversation set and differing category exclusions limit cross-report comparisons.
LongBench v2 · 2 · 503 questions · Source guide
Can it reason over a very long document? Reading comprehension across documents, conversations, code repositories, and structured data.
- Measure
- Multiple-choice accuracy, with short, medium, and long context breakdowns.
- Comparison rule
- Match chain-of-thought setting, input truncation, length bucket, and model context limit.
- Long-context reading does not measure persistent memory between sessions.
- Advertised context capacity does not guarantee useful recall at that length.
BrowseComp · 2025 · 1,266 questions · Source guide
Can it find a hard-to-locate fact on the web? Persistent search and multi-step browsing, scored through short factual answers.
- Measure
- Accuracy on web questions that require extensive information seeking.
- Comparison rule
- Match search tools, context management, agent count, and retry policy; label developer-reported scores.
- It does not directly evaluate the quality of a long research report.
- Live search results and different harnesses make cross-release scores difficult to compare.
BrowseComp-Plus · ACL 2026 · fixed corpus · Source guide
How good is the research agent when search access is controlled? A reproducible information-retrieval setting for difficult research questions.
- Measure
- Answer accuracy and retrieval effectiveness over a released document collection.
- Comparison rule
- Keep corpus revision, retriever, retrieval budget, and agent configuration fixed.
- A fixed corpus cannot represent the freshness or changing access conditions of the live web.
- BrowseComp-Plus scores cannot be substituted for BrowseComp scores.
Explore BrowseComp-Plus · Waterloo / CSIRO / collaborators · Methodology
FrontierMath · Tiers 1–3 / Tier 4 · v2 · Source guide
Can it solve a difficult research-level math problem? Expert-authored mathematical problems with automatically checkable answers.
- Measure
- Solved-problem rate under the evaluator's tools and compute budget.
- Comparison rule
- Keep version, difficulty tier, holdout set, and tool budget explicit.
- Tier 4 and Tiers 1–3 answer different difficulty questions.
- OpenAI funded the original benchmark and has access to some problems; Epoch describes the held-out subsets.
AstaBench · Scientific research suite · Source guide
Can it carry out the steps of scientific research? Evidence spanning literature work, code execution, data analysis, and discovery.
- Measure
- Task-specific research outcomes across a suite of scientific evaluations.
- Comparison rule
- Select a named sub-benchmark and matching tools; retain agent configuration and uncertainty.
- Not every submitted agent supports every task family.
- A suite average can hide a strong specialty and a missing capability.
SciCode · 2024 · main problems · Source guide
Can it translate scientific knowledge into working code? Coding problems drawn from numerical methods, simulations, and scientific calculations.
- Measure
- Main-problem or subproblem correctness against scientific test cases.
- Comparison rule
- Keep background-information setting, subproblem assistance, and model-generated versus gold earlier steps separate.
- An assisted subproblem score is not an end-to-end scientific workflow score.
- This adds more value inside the science category than as another homepage coding total.
CritPt · Research-level physics · Source guide
Can it reason through an unfamiliar physics research problem? Challenging physics tasks that test reasoning beyond standard academic exams.
- Measure
- Accuracy on the evaluator's research-level physics problems.
- Comparison rule
- Keep the evaluation revision, tools, reasoning budget, and grading protocol fixed.
- A specialist physics score should not stand in for general scientific usefulness.
- Difficulty and domain coverage differ from HLE and Terminal-Bench-Science.
Explore CritPt · Artificial Analysis / CritPt authors · Methodology
LiveBench · 2026-06-25 · Source guide
How does it handle fresh, objectively scored tasks? A periodically refreshed collection covering reasoning, language, data, instruction following, and coding.
- Measure
- Objective task scores and category results on a named release.
- Comparison rule
- Compare only the same release and task categories; refreshes change the exam.
- Its broad score overlaps with other general-purpose indices.
- Freshness reduces contamination risk; it does not establish its absence.
LiveCodeBench · Release v6 · dated windows · Source guide
Can it solve a new programming problem? Competitive-programming problems published over time, with executable tests.
- Measure
- Code-generation correctness on a specified problem-date window and scenario.
- Comparison rule
- Match release, start/end dates, scenario, sampling count, and execution budget.
- Algorithmic problem solving does not establish repository-editing skill.
- Different date windows are different cohorts, even if both are called v6.
Humanity’s Last Exam · Classic · 2025-04-03 final set · Source guide
Can it answer difficult questions across expert fields? An academic breadth check spanning mathematics, science, and the humanities, with text and image questions.
- Measure
- Answer accuracy on the finalized 2,500-question classic set; calibration is a separate measure.
- Comparison rule
- Pin the dataset revision, full versus text-only set, tools, reasoning effort, and answer judge. HLE-Rolling is a different exam.
- Closed-ended expert questions do not measure open-ended discovery or professional work.
- Tool-assisted and no-tool scores cannot be pooled.
- The dataset is gated; this guide links to it without redistributing its questions.
Explore Humanity’s Last Exam · Center for AI Safety / Scale AI · Methodology
GPQA Diamond · Diamond · 198 questions · Source guide
Can it reason through an expert science question? A compact multiple-choice test in biology, chemistry, and physics, selected through expert and non-expert review.
- Measure
- Multiple-choice answer accuracy on the 198-question Diamond subset; random choice has a 25% baseline.
- Comparison rule
- Keep Diamond separate from Main and Extended. Match prompts, answer shuffling, tools, reasoning budget, and sampling policy.
- A small fixed exam does not establish scientific discovery or laboratory competence.
- Near-ceiling scores and small gaps require uncertainty, not a confident rank order.
- Best-of-many success is not single-attempt accuracy.
SWE-bench Verified · Verified · 500 tasks · Source guide
Can it repair an issue in an existing repository? Real repository issues with executable tests, useful for understanding a coding agent's patching ability.
- Measure
- Percentage of the 500 human-filtered task instances resolved.
- Comparison rule
- Use one task revision, agent, budget, and attempt policy. The Bash Only view controls the agent; the full board mixes systems.
- mini-SWE-agent 1.x and 2.x change action handling and sampling and are not automatically comparable.
- Public task exposure and test quality limit conclusions about new, unseen work.
- A repair score does not measure an entire software development workflow.
SWE-bench Pro · Public · 731 tasks · Source guide
Can it make a larger change in a complex codebase? Longer software-engineering tasks across public application and developer-tool repositories.
- Measure
- Resolve rate: the percentage of public tasks whose patches pass the required new and regression tests.
- Comparison rule
- Keep Public, Private, and Held-out sets separate. Match dataset revision, agent harness, turn limit, cost cap, and attempts.
- The owner board mixes harnesses and capped versus uncapped runs; its rows are not one controlled cohort.
- Public repository licensing is not evidence that models have never seen the code.
- Task-quality disputes make this supporting evidence, not a universal replacement for Verified or Terminal-Bench.
CursorBench · 3.2 · Source guide
Which model handles the kind of work done inside Cursor? Ambiguous multi-file tasks drawn from Cursor usage, including instruction following and advanced tool use.
- Measure
- Cursor's task-correctness score in percent, alongside average API-priced cost per task, tokens, and steps.
- Comparison rule
- Compare only version 3.2 with its Cursor agent setup and exact model effort. Preserve the pricing revision for cost comparisons.
- This is a vendor-owned internal evaluation, not an independent cross-product ranking.
- Private tasks and agentic grading limit external reproduction.
- The task set changed from 3.1; small score gaps may reflect evaluation variance.
Explore CursorBench · Cursor · vendor-reported · Methodology
GDPval · 2025 · original evaluation · Source guide
Can it deliver work an experienced professional would accept? Occupational tasks with reference files and finished deliverables, including documents, presentations, and spreadsheets.
- Measure
- Quality of completed work compared with expert deliverables across 44 occupations; the public gold set contains 220 tasks.
- Comparison rule
- Name the full or gold task set, grading protocol, scaffolding, and whether ties count toward the reported win rate.
- The original evaluation and Artificial Analysis's GDPval-AA use different evaluation protocols.
- Producing one deliverable is not the same as doing an entire job or handling its organizational context.
- Developer-reported results need to retain their exact human or model-judge protocol.
GDPval-AA · v2 · Source guide
Which tool-using model produces the strongest professional deliverable? Artificial Analysis evaluates GDPval work products in its Stirrup agent environment, then compares the outputs head to head.
- Measure
- Pairwise Elo rating anchored to a human-expert baseline of 1,000, with source-reported uncertainty.
- Comparison rule
- Keep v2, the Stirrup environment, reasoning effort, judge panel, and rating pool together. Elo is not percent correct.
- The v2 environment, turn limit, and panel of judges differ from v1.
- Model-judged preferences are not a direct measurement of business value or worker replacement.
- Per-task cost must not be mixed with full-evaluation spend.
OSWorld 2.0 · osworld-v2-2026.08.08 · Source guide
Can an agent finish a workflow across desktop and web apps? Long computer-use tasks with verifiable outcomes, not just recognizing a button in a screenshot.
- Measure
- Task completion and partial reward across the pinned 108-workflow release, reported separately.
- Comparison rule
- Pin code, tasks, assets, website, provider image, step budget, and input/action interface before comparing systems.
- OSWorld-Verified and OSWorld 2.0 are different task cohorts.
- Partial progress is not a completed workflow.
- Gated environment assets and long runs affect reproducibility; the release manifest identifies the required versions.
τ³-bench · 3 · v1.0.1 grading · Source guide
Can a service agent solve the issue while following policy? Simulated customer-service conversations combine tool actions, user coordination, and domain rules; newer tracks add knowledge retrieval and voice.
- Measure
- Task success and repeated-trial reliability within a named domain and communication mode.
- Comparison rule
- Match domain, task split, user simulator, trials, and text or voice mode. pass^k consistency is not pass@k best-of-k success.
- The repository retains the tau2-bench name while the current suite is τ³-bench.
- Banking-knowledge results before v1.0.1 are not comparable with the corrected grading.
- Simulated service interactions do not cover every live customer or organizational policy.
WISE Verified · Verified · Qwen3.5-35B-A3B · 29 charted results
Can it draw what a prompt implies? Image generation that needs world knowledge: culture, time, space, biology, physics, and chemistry.
- Measure
- Weighted knowledge-consistency score, 0–1; higher is better.
- Comparison rule
- Only the Verified prompts and Qwen3.5-35B-A3B judge. Keep agent and chain-of-thought configurations named.
- Evaluation cohort
- Same 1,000 Verified prompts and Qwen3.5-35B-A3B judge. Chain-of-thought and agent systems remain separate configurations. This measures world-knowledge consistency, not aesthetic preference.
- Score
- Knowledge consistency (score); higher is better.
- Source retrieved
- Source revision
sha256:1b06a1e2697244ff1f29417662bc34b5708587c6bf335e0de65c80d984972263
- Not an aesthetic preference or image-editing test.
- The Verified prompts and judge changed in 2026; legacy WISE scores are not comparable.
- Published research cohort, not a complete inventory of today's image models.
Explore WISE Verified · WISE benchmark team · Methodology · Download this dataset
GEditBench · 2 · 16 charted results
Can it make an edit without breaking the rest? Instruction following, visual quality, and preservation of the original image across 23 editing tasks.
- Measure
- Overall pairwise Elo; higher is better. Publisher bootstrap intervals included.
- Comparison rule
- GPT-4o judges instruction and quality; PVC-Judge scores consistency. Compare within this fixed evaluation pool.
- Evaluation cohort
- GPT-4o judges instruction following and quality; PVC-Judge scores preservation. Elo and bootstrap intervals belong only to this 16-model research pool. Evaluated sample counts vary by model.
- Score
- Editing overall (Elo); higher is better.
- Source retrieved
- Source revision
sha256:87a97fa869d8d879e2f834520a109632e41c7f8e43724aabae6bb13882f838cc
- Automated, human-aligned judging is not direct human voting.
- This is the 2026 paper cohort; the dated API names are retained.
- Safety refusals and failures leave some models with fewer evaluated samples.
- Elo is pool-relative, not a percentage; overlapping intervals do not establish a winner.
Explore GEditBench · GEditBench v2 team · Methodology · Download this dataset
VideoPhy · 2 · human evaluation · 7 charted results
Does the action obey basic physics? Human reviewers check whether a generated video both follows the prompt and respects physical commonsense.
- Measure
- Joint semantic and physical adherence, %; higher is better.
- Comparison rule
- Human-evaluated All subset only. Hard, physical-activity, and object-interaction scores remain separate details.
- Evaluation cohort
- Percentage satisfying both semantic adherence and physical commonsense in the All subset. These 2025 human evaluations are a historical diagnostic, not a current video-model buying guide.
- Score
- Prompt + physical adherence (%); higher is better.
- Source retrieved
- Source revision
sha256:bad716a62a9247323932ca8ea305feb4a2f373ebe01233beece8e3dbea4fbd34
- A 2025 research cohort, not a current ranking of video generators.
- Physical plausibility is different from cinematic appeal.
- Automatic VideoPhy2-eval scores must not be mixed with these human judgments.
Explore VideoPhy · VideoPhy2 team · Methodology · Download this dataset
OmniDocBench · 1.6_full · 35 charted results
Can it turn a difficult document into usable text? Document extraction across text, formulas, tables, and reading order, comparing specialist pipelines with general vision models.
- Measure
- Publisher overall score, 0–100; higher is better. Component error rates retain their native direction.
- Comparison rule
- Use the exact v1.6_full result table; do not pool v1.0, v1.5, or other document datasets.
- Evaluation cohort
- The publisher's v1.6_full table is retained even though the repository now advertises v1.7. Overall combines text, formula, and table performance. Component error metrics are lower-is-better; CDM and TEDS are higher-is-better.
- Score
- Document extraction overall (/ 100); higher is better.
- Source retrieved
- Source revision
sha256:a799f3dabee0d8e2be9f990fc325c3fefadb8d7f8d0cf56148406547d495a3ad
- The repository advertises v1.7 but still labels this model table v1.6_full; its table label is preserved.
- A parsing system and a general vision model are different deployment choices.
- No matched latency or cost measurements are provided in this extract.
Explore OmniDocBench · OpenDataLab · Methodology · Download this dataset
WorldScore · Static · 2025 author cohort · 19 charted results
Can a generated world stay coherent as the camera moves? A common scene-generation protocol compares camera control, scene consistency, and quality across video, 3D, and 4D systems.
- Measure
- WorldScore-Static, 0–100; higher is better.
- Comparison rule
- Only systems sampled and evaluated by the WorldScore authors on March 30, 2025. Input and system types stay visible.
- Evaluation cohort
- All 19 systems were sampled and evaluated by the WorldScore authors on March 30, 2025. Static scene quality and control are compared in one protocol; this does not rank today's interactive world simulators. Newer self-evaluated submissions are excluded.
- Score
- WorldScore-Static (/ 100); higher is better.
- Source retrieved
- Source observation date
- Source revision
sha256:a1c51fa9e6f1d4970f58293f4f648dda890a68d987f48b21ab8e61e0542adff3
- Historical research comparison; newer model-team submissions are not in this chart.
- Static and Dynamic are distinct scores; this chart does not measure interactive control latency or physical simulation accuracy.
- A strong generated video is not evidence of an action-conditioned world simulator.
Explore WorldScore · WorldScore team · Methodology · Download this dataset
GenEval2 · 2025-12 · Source guide
Are the objects, attributes, and relationships correct? Checks the detailed compositional content of generated images, beyond whether an image looks good.
- Measure
- Soft-TIFA compositional alignment; higher is better.
- Comparison rule
- Keep arithmetic and geometric aggregation separate; GenEval2 is not the original GenEval score.
- Judge and prompt-processing settings affect results.
- Use the same benchmark release and scoring configuration when comparing results.
DPG-Bench · 2024-03 · Source guide
Does a detailed image prompt survive intact? Breaks dense text-to-image prompts into questions about the requested content.
- Measure
- Dense-prompt alignment score; higher is better.
- Comparison rule
- Require identical prompts, question dependencies, VQA evaluator, and rewriting policy.
- Prompt rewriting can change what is being tested.
- A legacy compositional test, not evidence of current aesthetic preference or editing quality.
T2I-CompBench++ · ++ · 2024 suite · Source guide
Can it bind the right attribute to the right object? Separates color, shape, texture, relationships, counting, and complex scene composition.
- Measure
- Dimension-specific composition scores; higher is better.
- Comparison rule
- Use the same ++ task split and evaluator per dimension; avoid a made-up average of unlike evaluators.
- The original suite and ++ extension have different task coverage.
- Detector and visual-judge errors can look like generation failures.
VBench · 2.0 · 2025 · Source guide
Where does a video generator break down? A diagnostic view of human fidelity, creativity, controllability, physics, and commonsense across 18 dimensions.
- Measure
- Dimension and five-aspect aggregate scores; higher is better.
- Comparison rule
- Keep VBench 2.0 separate from VBench 1.0, VBench++, and image-to-video tracks.
- The overall score gives equal weight to five aspects, not to every dimension.
- Submission settings and evaluator revisions affect comparability.
- Automatic diagnostics do not replace viewer preference.
WorldModelBench · 2025 · Source guide
Does a predicted world follow the instruction and physical rules? Image-conditioned world generation judged across everyday, simulated, and embodied environments.
- Measure
- Instruction, commonsense, and physical-adherence scores; higher is better.
- Comparison rule
- Hold the input images, environment domains, and world-model evaluator fixed.
- A small research suite is not an interactive-agent deployment test.
- The judge and evaluated model cohort matter; a newer demo is not a benchmark result.
Explore WorldModelBench · WorldModelBench team · Methodology
Open ASR Leaderboard · Public evaluation tracks · Source guide
Which system transcribes speech accurately and quickly? Speech recognition across datasets and languages, with accuracy and inference speed reported separately.
- Measure
- Word error rate: lower is better. Real-time factor speedup: higher is better.
- Comparison rule
- Match language, dataset, chunking, decoding settings, and hardware before comparing speed.
- Short-form, long-form, and multilingual tracks are different cohorts.
- GPU throughput is not end-to-end API latency.
- Hardware, chunking, and decoding settings must match for a fair efficiency comparison.
Explore Open ASR Leaderboard · Hugging Face audio team · Methodology
SEED-TTS-Eval · 2024 · Source guide
Can it say the right words in the requested voice? Zero-shot speech synthesis evaluated for intelligibility and similarity to a reference speaker in English and Mandarin.
- Measure
- Word error rate: lower is better. Speaker similarity: higher is better.
- Comparison rule
- Keep language, ASR scorer, speaker encoder, and voice-conditioning protocol fixed.
- Speaker similarity and correct words do not measure expressive naturalness.
- Voice-cloning evaluation is distinct from choosing a ready-made production voice.
Voice Arena · TTS v1 · Source guide
Which synthetic voice do listeners prefer? Pairwise listener preference for speech synthesis, separated by language.
- Measure
- Language-specific preference rating; higher is better.
- Comparison rule
- Compare within the same language and voice setup; retain uncertainty and vote counts.
- Different language pools are not on one universal scale.
- Listener preference does not establish word-perfect transcription, speaker cloning, or conversational latency.
MMAU-Pro · 2025-08 · Source guide
Can it reason about what it hears? Audio understanding across speech, environmental sounds, and music, including long and multiple recordings.
- Measure
- Task-specific audio reasoning accuracy; higher is better.
- Comparison rule
- Keep multiple-choice, open-ended, and instruction-following evaluation modes distinct.
- Not a speech-generation or music-generation preference test.
- Transcription-only systems do not receive the same sensory input as audio-native models.
MMMU-Pro · 2024 suite · Source guide
Can it reason from a diagram, not just read its text? Expert-level multimodal problems designed to reduce text-only shortcuts.
- Measure
- Accuracy, %; higher is better.
- Comparison rule
- Keep Standard 10-option and Vision settings named; original MMMU and MMMU-Pro are not interchangeable.
- Prompting, resolution, tools, and reasoning effort affect the result.
- Academic question answering is not document parsing or visual design quality.
Video-MME · 2024 / CVPR 2025 · Source guide
Can it understand a long video? Video question answering at short, medium, and long durations.
- Measure
- Question-answer accuracy, %; higher is better.
- Comparison rule
- Keep with-subtitle and without-subtitle tracks separate; record frame sampling and audio access.
- Extra frames, subtitles, and audio change the information supplied to the model.
- Understanding video is different from generating it.
Image Arena · Live · text-to-image · Source guide
Which generated image do people prefer? Blind human preference provides an aesthetic and overall-utility view alongside diagnostic image benchmarks.
- Measure
- Pairwise Elo with confidence intervals; higher is better.
- Comparison rule
- Text-to-image and editing have separate opponent pools; retain model settings, votes, and rating uncertainty.
- Preference is not a factuality or exact-composition guarantee.
- Use the publisher’s current pool, vote counts, and intervals when comparing its scores.
Video Arena · Live · generation tracks · Source guide
Which video looks best to viewers? Human preference for generated video, with separate input and audio tracks.
- Measure
- Pairwise Elo with confidence intervals; higher is better.
- Comparison rule
- Keep text-to-video, image-to-video, and with-audio pools separate; match resolution, duration, and frame rate for cost.
- A silent-video score does not describe audio quality.
- Use the publisher’s matching input, duration, resolution, and audio track when comparing results.
Open ASR · meeting transcription · AMI-Cleaned · English test · 2026-09-04 · 10 charted results
Which open-weight speech models make fewer errors in English meetings? Ten selected configurations on the same cleaned meeting-transcription test. These are publisher-reported results with different inference pipelines.
- Measure
- Word error rate (WER), %; lower is better. Counts substituted, missing, and extra words.
- Comparison rule
- Compare only the AMI-Cleaned English test in this September 4, 2026 snapshot. Keep the model version and publisher scoring protocol fixed.
- Evaluation cohort
- Same AMI-Cleaned English test and publisher scoring protocol. Inference pipelines differ; exact per-run configurations are not published. Lower word-error rate is better. This is not a speed, speaker-attribution, or overall audio-quality ranking.
- Score
- Word error rate (% WER); lower is better.
- Source retrieved
- Source observation date
- Source revision
ba5712d5ace8f785fa0daae1aecea8561ecd87c9
- A selected open-weight comparison, not the full leaderboard or a top-ten list.
- Inference pipelines differ. The published rows do not identify exact run dates, decoding settings, checkpoint revisions, or execution records.
- Short speech clips do not test whole-meeting speaker attribution, punctuation quality, other languages, or live response time.
- No uncertainty is published; small differences do not establish a reliable winner. WER can exceed 100% when extra words are inserted.
Explore Open ASR · meeting transcription · Hugging Face Open ASR Leaderboard · Methodology · Download this dataset
Terminal-Bench 4 coding standard
Terminal-Bench 4.0.0 is the site’s standard agentic terminal-engineering benchmark. Explore it in the benchmark library. The checked snapshot contains 13 configurations from the official Harbor Framework submissions at commit c7878bd, committed on , with 66 tasks and 5 trials per task. aicharts retrieved this owner snapshot on .
This standalone Terminal-Bench 4 owner cohort remains separate from the Terminal-Bench 4 component reported inside the Artificial Analysis Coding Agent Index. Every standalone TB4 row retains its model, agent, agent version, effort, accuracy, 95% confidence interval, trials, cost, tokens, duration, and pinned source files.
Download Terminal-Bench 4 JSONTerminal-Bench-Science 0.1
The checked scientific-workflow snapshot published by Terminal-Bench-Science and Harbor Framework contains 17 owner-published system configurations across 70 tasks and 3 trials per task. It is pinned to the exact v0.1.0 release commit f81afac. The release has the persistent citation https://doi.org/10.5281/zenodo.22110254, and the owner leaderboard was updated on . aicharts retrieved it on .
Every row keeps the named model, harness, reasoning effort, resolution rate, binomial standard error, trial count, evaluation cost, token use, and owner-published source link. Terminal-Bench-Science remains separate from general terminal engineering and does not feed a composite score. Owner-published aggregate and per-domain cost fields are retained independently and are not forced to reconcile.
Download Terminal-Bench-Science 0.1 JSONCurrent Intelligence efficiency · v4.3.2
The homepage’s Pareto chart uses Artificial Analysis Intelligence Index v4.3.2. Its output-token and cost views compare the identical 97-configuration positive-cost cohort from 101 complete score-and-output records. Output tokens include answer and reasoning; task cost also includes input and cache traffic.
This 10-evaluation version weights agents 30%, coding 20%, scientific reasoning 20%, and general capability 30%. The source is checked every four hours. Version, evaluation roster, source identities, native measures, and retention must pass validation before an update is published. Retrieved .
Download current Intelligence v4.3.2 JSON · Full v4.3.2 comparison rules and limitations. Keep these results separate from the frozen v4.1.1 snapshot below; its evaluation roster and weights differ.
Historical Intelligence v4.1.1 · frozen snapshot
This retained historical dataset pairs the owner-published Artificial Analysis Intelligence Index v4.1.1 score with weighted output tokens and cost per Intelligence Index task. It is not the current homepage dataset. The frozen snapshot retains 135 measured score-and-output records from 643 source records. Its historical matched-resource comparison uses the same 127-record cohort with positive comparable cost.
The Index combines 9 evaluations. Its four owner-defined category weights and constituent evaluation weights are:
- Agents · 34%: GDPval-AA v2 · 20%; τ³-Banking · 14%.
- Coding · 24%: Terminal-Bench v2.1 · 16%; SciCode · 8%.
- Scientific Reasoning · 24%: Humanity's Last Exam · 12%; GPQA Diamond · 6%; CritPt · 6%.
- General · 18%: AA-LCR · 6%; AA-Omniscience · 12%.
Output tokens here mean answer plus reasoning tokens only, weighted by each evaluation's Index weight and divided by its task count. They are not the coding-agent chart's total tokens, which also include input traffic. Cost is the owner's weighted per-task sum of available input, cache, reasoning, and answer/output components. A source row with a complete cost breakdown but a reported zero total is stored as unavailable, never converted into a free-model value. Rows with incomplete cost are excluded. Complete zero-total rows remain in the JSON but are omitted from the historical matched-resource cohort.
The comparable cohort follows the checked rule: current, non-estimated model configurations with a finite 0–100 Intelligence Index, finite positive per-task output-token total, and a complete finite nonnegative per-task cost breakdown; source cost totals at or below zero normalize to null. In a Pareto frontier for that stored cohort, a frontier point is not dominated by another record with an equal-or-higher Intelligence score and equal-or-lower output-token or positive-cost value. Artificial Analysis publishes the measurements; the frontier classification is aicharts analysis.
This v4.1.1 snapshot is frozen and is no longer refreshed by automation. The current v4.3.2 dataset above has a separate versioned source contract and download. aicharts retrieved this historical snapshot on .
Read the owner's Index methodology, model leaderboard, and terms of use. Citation: Artificial Analysis (2025). LLM benchmarks dataset. https://artificialanalysis.ai.
Download historical v4.1.1 JSONArtificial Analysis coding-agent source and refresh
The source is the public Coding Agent Index v1.5 Artificial Analysis coding-agents comparison, composed of DeepSWE v1.1, Terminal-Bench 4, and SWE-Atlas-QnA. aicharts retrieved this snapshot on . The site checks for a new source snapshot daily. The displayed retrieval time changes only when a validated snapshot is stored.
The most recent retained model, variant, or material benchmark change was detected on . This meaningful-update time is separate from the daily retrieval check.
The checked dataset contains 20 model-agent configurations across 19 models, 8 agent harnesses, and 10 model providers. The chart and the JSON download use this same checked snapshot.
Benchmark definitions
Each benchmark is shown on the 0–100 scale stored in the snapshot. The metrics evaluate different tasks and should be interpreted separately.
AA Index
Overall performance across code changes, terminal work, and repository understanding.
DeepSWE v1.1
Long-horizon software engineering tasks scored with automated code verification.
Terminal-Bench 4
Agentic terminal-use tasks scored with automated test-suite verification.
SWE-Atlas-QnA
Repository-understanding questions scored with a strict resolve verifier.
Current leaders
These are the highest available scores in the retrieved snapshot, one row per benchmark. They are observations of the named model, agent harness, and effort setting rather than general model ranks.
| Benchmark | Model | Agent | Provider | Setting | Score |
|---|---|---|---|---|---|
| AA Index | Opus 5.5 | Claude Code | Anthropic | max | 66.0 |
| DeepSWE v1.1 | Muse Spark 1.3 | Muse Code | Meta | xhigh | 73.2 |
| Terminal-Bench 4 | Opus 5.5 | Claude Code | Anthropic | max | 63.1 |
| SWE-Atlas-QnA | Opus 5.5 | Claude Code | Anthropic | max | 66.4 |
For AA Index versus mean API cost, including the cost/performance frontier, see highest AA Index and lowest cost pick different agents. For whether classified open-weight rows sit with those leaders, see open models closed SemiAnalysis composites, not this table. For how a cheaper model changed one daily news page, see GPT-5.6 Luna made one daily news page cost about $0.10. For what a 30% Terminal-Bench-Science result measures, and how cost and token use change the comparison, see What Terminal-Bench-Science’s 30% result measures. For why a public-suite high score still needs a holdout, see why a coding-agent high score still needs a holdout.
All configurations
Every model-agent configuration in the retrieved snapshot, with AA Index, component scores, and mean API cost per task. Missing values are stored as empty in the source and shown as a dash.
| Model | Agent | Provider | Setting | AA Index | DeepSWE v1.1 | Terminal-Bench 4 | SWE-Atlas-QnA | Cost |
|---|---|---|---|---|---|---|---|---|
| Opus 5.5 | Claude Code | Anthropic | max | 66.0 | 68.4 | 63.1 | 66.4 | $13.04 |
| Fable 5.1 (with fallback) | Claude Code | Anthropic | max | 62.2 | 64.3 | 57.6 | 64.8 | $12.39 |
| Claude Fable 5.1 XHigh + SWE-2 Medium | Devin Fusion CLI | Cognition | default | 61.7 | 63.1 | 56.1 | 65.9 | $7.90 |
| GPT-6 Astra | Codex | OpenAI | max | 61.6 | 67.6 | 55.6 | 61.8 | $7.47 |
| Opus 5 | Claude Code | Anthropic | max | 59.7 | 62.5 | 54.5 | 62.1 | $10.79 |
| GPT-6 Astra XHigh + SWE-2 Medium | Devin Fusion CLI | Cognition | default | 58.9 | 67.3 | 50.0 | 59.4 | $4.54 |
| GPT-6 Sol | Codex | OpenAI | max | 56.7 | 69.0 | 43.4 | 57.5 | $2.99 |
| Grok 4.7 | Grok Build | xAI | xhigh | 56.3 | 72.6 | 33.3 | 62.9 | $8.82 |
| GPT-5.6 Sol | Codex | OpenAI | max | 54.6 | 72.3 | 37.4 | 54.0 | $6.35 |
| Muse Spark 1.3 | Muse Code | Meta | max | 54.3 | 71.7 | 31.8 | 59.4 | $3.98 |
| GLM-5.3 | Opencode | Z.ai | default | 53.6 | 61.4 | 39.9 | 59.4 | $4.24 |
| Kimi K3 | Kimi Code CLI | Moonshot AI | default | 51.9 | 68.4 | 21.2 | 66.1 | $5.05 |
| Muse Spark 1.3 | Muse Code | Meta | xhigh | 48.3 | 73.2 | 17.2 | 54.6 | $3.47 |
| Grok 4.6 | Grok Build | xAI | xhigh | 47.0 | 64.9 | 17.7 | 58.3 | $3.57 |
| Qwen3.8 Max | Claude Code | Alibaba Cloud | default | 43.3 | 51.0 | 16.7 | 62.1 | $3.48 |
| GPT-5.6 Luna | Codex | OpenAI | max | 43.2 | 66.4 | 14.6 | 48.7 | $0.438 |
| DeepSeek V4 Pro 0813 | Codex | DeepSeek | max | 43.1 | 57.2 | 10.1 | 61.8 | $0.238 |
| Gemini 3.8 Flash | Antigravity SDK | high | 41.9 | 65.8 | 14.6 | 45.2 | $2.47 | |
| GPT-6 Luna | Codex | OpenAI | max | 41.1 | 63.7 | 15.2 | 44.4 | $0.176 |
| DeepSeek V4 Flash 0731 | Codex | DeepSeek | max | 38.7 | 54.3 | 10.6 | 51.3 | $0.085 |
Normalization method
The refresh job reads the source page's public data payload, validates every source row, and maps it into a versioned owned schema. Benchmark reward proportions are represented as 0–100 scores. Mean task cost stays in US dollars, mean active wall time stays in seconds, and mean total token use stays as a token count.
Provider identifiers, model effort settings, stable series keys, and sort order are normalized for the chart. The refresh is rejected when duplicate records, major row loss, stable-key loss, or substantial metric-coverage regressions are detected. aicharts does not recalculate the source benchmark outcomes.
Limitations
- Artificial Analysis defines and operates the upstream evaluations. aicharts is an independent visualization and is not affiliated with Artificial Analysis or the listed providers.
- The Intelligence efficiency view is an owner-defined, primarily English-language aggregate. Its category weights emphasize agentic tasks, and it does not establish performance for every use case.
- Scores depend on the named model, agent harness, effort setting, task set, and evaluation version. They do not establish results for every software repository or production workflow.
- Cost, duration, and token values are task-level means from the source evaluation. They are not price or latency guarantees.
- The current v4.3.2 Intelligence snapshot is checked every four hours and the coding-agent snapshot daily; neither is a real-time mirror. The historical v4.1.1 snapshot is frozen. Use the relevant version and retrieval timestamp when citing a value.