Skip to benchmark notes

Terminal-Bench-Science: 30% is not a product win

Scientists, not vendors, set the bar on Terminal-Bench-Science 0.1. The peak 30% resolution is remaining work, not a shipping product. Cost and token Pareto is the useful comparison.

Paper-cutout hands hold a high charcoal measuring beam while a compact cobalt apparatus reaches only partway, with coral and mint weights on a low tray.
Scientists set the evaluation bar; a peak score without cost and token trade-offs is not a product win. The illustration is not a data plot. AI Charts editorial illustration · Atet with GPT Image 2

Terminal-Bench-Science 0.1 is a Stanford-led benchmark of AI agents on scientific research workflows. Steven Dillmann, writing the announcement for the Terminal-Bench-Science team, says “Scientists, not model developers or data vendors, set the bar for scientific capability in AI.” The suite is built with the Terminal-Bench and Harbor team and with domain experts. Of 920 proposals, only 70 tasks survived domain, technical, and bar-raiser review. “The strongest model evaluated, Claude Opus 5, achieves a 30% resolution rate on Terminal-Bench-Science 0.1.”

That 30% cell is the headline. It is not a product win. A research assistant that fails most of the workflows scientists chose to measure is still a research object. The useful comparison on AI Charts is the one the suite already reports next to resolution: cost and token Pareto. The coding-agent comparison chart already plots those axes for a different, software-engineering snapshot.

The Hraness reading note of that announcement, saved 2026-08-29, is a dated digest. It is not a substitute for the primary page.

Scientists set the evaluation bar

Tasks come from researchers’ own work across the life, physical, Earth, mathematical, and engineering sciences. They are not textbook questions or standardized exercises. Contributors propose workflows. Reviewers approve a subset for implementation. A later review asks whether each task is objectively verifiable, hard for current frontier agents, and scientifically real.

“Of 920 proposals, 464 were approved for implementation and 386 pull requests were opened, but only 70 tasks made it into Terminal-Bench-Science 0.1.” Dillmann presents that selectivity as evidence that it is hard to write tasks that are scientifically interesting, hard for frontier agents, and specified well enough to grade. The accepted set covers scientific data analysis, statistical inference, simulation, optimization, theorem proving, image reconstruction, signal processing, inverse problems, sensor calibration, model fitting, classification, and scientific machine learning.

Hraness defines an agent harness as software that gives a model a place to work: it injects instructions, offers tools, runs an assess-act-reassess loop, and translates across model APIs. The Science leaderboard names that layer. Claude Code, Codex, and Grok Build are part of each observation. A model name without the harness is an incomplete citation here too.

A 30% peak is remaining work

Each evaluated model ran three independent trials per task across all 70 tasks. Claude Opus 5 with Claude Code leads at 30.0%. GPT-5.6 Sol with Codex is next at 22.4%, then Claude Fable 5 with Claude Code at 21.4%. Claude Opus 4.8 sits at 10.5%. Several other named systems resolve less than 11%. GLM 5.3 with Claude Code is the strongest open model at 8.1%. GPT-5.6 Luna with Codex is last at 3.3%.

Terminal-Bench-Science 0.1 resolution rates from the announcement
ModelAgent harnessResolution
Claude Opus 5Claude Code30.0%
GPT-5.6 SolCodex22.4%
Claude Fable 5Claude Code21.4%
Claude Opus 4.8Claude Code10.5%
GPT-5.6 TerraCodex8.6%
GLM 5.3Claude Code8.1%
Kimi K3Claude Code7.1%
Grok 4.6Grok Build7.1%
GPT-5.6 LunaCodex3.3%

The suite is calibrated to sit more than 10 percentage points below Terminal-Bench 3.0 for every model evaluated on both. Reviewers rejected tasks that frontier agents already solve easily. A 30% resolution rate therefore means the leading configuration still failed most of the remaining, scientist-set workflows. That is a ceiling on current capability, not a claim that a lab can replace the scientist on the work the suite measures.

The announcement’s own goal is agents that execute demanding workflows so scientists can spend more time on questions, hypotheses, interpretation, and communication. A system that resolves three in ten of the accepted tasks does not yet occupy that role.

Cost and tokens split the useful ranking

“GPT-5.6 Sol matches Claude Fable 5's performance at less than a third of the cost ($4.2k vs $14.2k).” Claude Opus 5 reaches the highest resolution at $7.0k. On tokens, Claude Fable 5 matches GPT-5.6 Sol’s performance while using about a quarter fewer tokens (6.4B versus 8.4B). Only Kimi K3 and Claude Opus 5 appear on both the cost-resolution and token-resolution Pareto frontiers.

Cost and token facts Terminal-Bench-Science 0.1 reports beside resolution
ComparisonReported resultWhat it changes
GPT-5.6 Sol versus Claude Fable 5$4.2k versus $14.2kSimilar resolution at less than a third of the evaluation cost
Claude Opus 5 evaluation cost$7.0kHighest resolution at a higher total evaluation cost
Claude Fable 5 versus GPT-5.6 Sol tokens6.4B versus 8.4BSimilar resolution at about a quarter fewer tokens
Both Pareto frontiersKimi K3 and Claude Opus 5The only named systems on both the cost and token fronts

Peak resolution and the useful trade-off are different questions. A product team that can only cite the 30% cell has not yet answered which configuration is cheapest, or cheapest in tokens, for a given quality bar.

This chart already plots that trade-off

The current AI Charts coding-agent comparison plots AA Index, DeepSWE, Terminal-Bench v2.1, and SWE-Atlas-QnA against API cost, active time, or total token use. Those scores belong to a checked Artificial Analysis coding-agents snapshot. AI Charts retrieved it on Aug 30, 2026, 3:01 PM UTC. Terminal-Bench-Science 0.1 is not in that snapshot. The two Terminal-Bench names share a franchise and a review culture. They do not share a task set.

The snapshot’s highest stored Terminal-Bench v2.1 score is 91.0 for Gemini 3.7 Flash on Opencode at the high setting. That number is a software-engineering terminal-suite observation. It does not establish a Science resolution rate.

What transfers is the comparison shape. Terminal-Bench-Science publishes cost and token Pareto next to resolution because a peak score without those axes is an incomplete product signal. This host already charts that shape for the coding-agent snapshot. Open the coding-agent comparison chart to change axes. Read how cheaper AI models can make everyday products viable when the question is a frequent-use consumer feature rather than a scientist-set workflow.

A different question from cheaper everyday models

How cheaper AI models can make everyday products viable asks whether a lower-cost model can make a repeated product feature viable once it meets a written quality bar. This page asks whether a 30% peak on a scientist-set science suite is a shipping research assistant. Both notes treat cost as part of the result. They do not share a workload.

The small-models note can stay with one person’s news-page experiment and listed token prices. This note stays with Dillmann’s scientist-set bar, the 70-task funnel, and the cost and token frontiers the suite already publishes. Collapsing those questions into one “cheaper is better” headline would drop the bar, the harness, and the task set.

How to read Terminal-Bench-Science

Read Dillmann’s announcement for the task funnel, the 70-task coverage, the named resolution rates, the cost and token frontiers, and the living-benchmark roadmap. Read the Hraness reading note of that announcement for a dated digest. Read the Hraness harness definition when a row’s agent name needs a noun. The next release, 0.2, has a pull-request deadline of October 5, 2026. Tasks are versioned so Harbor can re-run trials.

The useful sentence is narrower than a leaderboard headline. Scientists set the bar. The leading named configuration resolves 30% of the accepted workflows. That remaining miss rate is the result. Cost and token Pareto is how AI Charts already asks the next product question.

Limits of this reading

  • Resolution rates, costs, and token totals belong to the named 0.1 release, models, harnesses, and three-trial protocol on the announcement. They can change in a later release.
  • Terminal-Bench-Science 0.1 is not in the checked Artificial Analysis coding-agent snapshot. A stored Terminal-Bench v2.1 score is a different suite.
  • Reported evaluation costs are totals across all 70 tasks. They are not a production invoice, a subscription price, or a per-query quote.
  • The suite is a living benchmark. Later releases will add, retire, and recalibrate tasks as the frontier moves.
  • Resolution varies by scientific domain. A suite score is not a domain score, and a domain lead is not a general research-assistant claim.
  • The claim that 30% is not a product win is AI Charts analysis of those reported rates. Cite Dillmann for the measurements.

Sources

  1. Terminal-Bench-Science 0.1Terminal-Bench-Science, 2026. Steven Dillmann’s announcement defines Terminal-Bench-Science 0.1, reports the 70-task funnel, named resolution rates, cost and token frontiers, and the living-benchmark roadmap.
  2. Hraness reading note: Terminal-Bench-Science 0.1Hraness, 2026. The Hraness reading note is a dated digest of the Terminal-Bench-Science 0.1 announcement, used here as a crawlable companion citation rather than a substitute for the original.
  3. Coding AgentsArtificial Analysis, 2026. The public coding-agents comparison is the upstream source of the checked AI Charts snapshot. Model names, agent harnesses, settings, AA Index scores, and mean API costs are Artificial Analysis measurements.

Reported results apply to the named source, workload, configuration, and observation date. They do not establish performance on every task or product.