Terminal-Bench-Science 0.1 is a Stanford-led benchmark of AI agents on scientific research workflows. Steven Dillmann, writing the announcement for the Terminal-Bench-Science team, says “Scientists, not model developers or data vendors, set the bar for scientific capability in AI.” The suite is built with the Terminal-Bench and Harbor team and with domain experts. Of 920 proposals, only 70 tasks survived domain, technical, and bar-raiser review. “The strongest model evaluated, Claude Opus 5, achieves a 30% resolution rate on Terminal-Bench-Science 0.1.”
That 30% cell is the headline. It is not a product win. A research assistant that fails most of the workflows scientists chose to measure is still a research object. The useful comparison on AI Charts is the one the suite already reports next to resolution: cost and token Pareto. The coding-agent comparison chart already plots those axes for a different, software-engineering snapshot.
The Hraness reading note of that announcement, saved 2026-08-29, is a dated digest. It is not a substitute for the primary page.
Scientists set the evaluation bar
Tasks come from researchers’ own work across the life, physical, Earth, mathematical, and engineering sciences. They are not textbook questions or standardized exercises. Contributors propose workflows. Reviewers approve a subset for implementation. A later review asks whether each task is objectively verifiable, hard for current frontier agents, and scientifically real.
“Of 920 proposals, 464 were approved for implementation and 386 pull requests were opened, but only 70 tasks made it into Terminal-Bench-Science 0.1.” Dillmann presents that selectivity as evidence that it is hard to write tasks that are scientifically interesting, hard for frontier agents, and specified well enough to grade. The accepted set covers scientific data analysis, statistical inference, simulation, optimization, theorem proving, image reconstruction, signal processing, inverse problems, sensor calibration, model fitting, classification, and scientific machine learning.
Hraness defines an agent harness as software that gives a model a place to work: it injects instructions, offers tools, runs an assess-act-reassess loop, and translates across model APIs. The Science leaderboard names that layer. Claude Code, Codex, and Grok Build are part of each observation. A model name without the harness is an incomplete citation here too.
A 30% peak is remaining work
Each evaluated model ran three independent trials per task across all 70 tasks. Claude Opus 5 with Claude Code leads at 30.0%. GPT-5.6 Sol with Codex is next at 22.4%, then Claude Fable 5 with Claude Code at 21.4%. Claude Opus 4.8 sits at 10.5%. Several other named systems resolve less than 11%. GLM 5.3 with Claude Code is the strongest open model at 8.1%. GPT-5.6 Luna with Codex is last at 3.3%.
| Model | Agent harness | Resolution |
|---|---|---|
| Claude Opus 5 | Claude Code | 30.0% |
| GPT-5.6 Sol | Codex | 22.4% |
| Claude Fable 5 | Claude Code | 21.4% |
| Claude Opus 4.8 | Claude Code | 10.5% |
| GPT-5.6 Terra | Codex | 8.6% |
| GLM 5.3 | Claude Code | 8.1% |
| Kimi K3 | Claude Code | 7.1% |
| Grok 4.6 | Grok Build | 7.1% |
| GPT-5.6 Luna | Codex | 3.3% |
The suite is calibrated to sit more than 10 percentage points below Terminal-Bench 3.0 for every model evaluated on both. Reviewers rejected tasks that frontier agents already solve easily. A 30% resolution rate therefore means the leading configuration still failed most of the remaining, scientist-set workflows. That is a ceiling on current capability, not a claim that a lab can replace the scientist on the work the suite measures.
The announcement’s own goal is agents that execute demanding workflows so scientists can spend more time on questions, hypotheses, interpretation, and communication. A system that resolves three in ten of the accepted tasks does not yet occupy that role.
Cost and tokens split the useful ranking
“GPT-5.6 Sol matches Claude Fable 5's performance at less than a third of the cost ($4.2k vs $14.2k).” Claude Opus 5 reaches the highest resolution at $7.0k. On tokens, Claude Fable 5 matches GPT-5.6 Sol’s performance while using about a quarter fewer tokens (6.4B versus 8.4B). Only Kimi K3 and Claude Opus 5 appear on both the cost-resolution and token-resolution Pareto frontiers.
| Comparison | Reported result | What it changes |
|---|---|---|
| GPT-5.6 Sol versus Claude Fable 5 | $4.2k versus $14.2k | Similar resolution at less than a third of the evaluation cost |
| Claude Opus 5 evaluation cost | $7.0k | Highest resolution at a higher total evaluation cost |
| Claude Fable 5 versus GPT-5.6 Sol tokens | 6.4B versus 8.4B | Similar resolution at about a quarter fewer tokens |
| Both Pareto frontiers | Kimi K3 and Claude Opus 5 | The only named systems on both the cost and token fronts |
Peak resolution and the useful trade-off are different questions. A product team that can only cite the 30% cell has not yet answered which configuration is cheapest, or cheapest in tokens, for a given quality bar.
This chart already plots that trade-off
The current AI Charts coding-agent comparison plots AA Index, DeepSWE, Terminal-Bench v2.1, and SWE-Atlas-QnA against API cost, active time, or total token use. Those scores belong to a checked Artificial Analysis coding-agents snapshot. AI Charts retrieved it on Aug 30, 2026, 3:01 PM UTC. Terminal-Bench-Science 0.1 is not in that snapshot. The two Terminal-Bench names share a franchise and a review culture. They do not share a task set.
The snapshot’s highest stored Terminal-Bench v2.1 score is 91.0 for Gemini 3.7 Flash on Opencode at the high setting. That number is a software-engineering terminal-suite observation. It does not establish a Science resolution rate.
What transfers is the comparison shape. Terminal-Bench-Science publishes cost and token Pareto next to resolution because a peak score without those axes is an incomplete product signal. This host already charts that shape for the coding-agent snapshot. Open the coding-agent comparison chart to change axes. Read how cheaper AI models can make everyday products viable when the question is a frequent-use consumer feature rather than a scientist-set workflow.
A different question from cheaper everyday models
How cheaper AI models can make everyday products viable asks whether a lower-cost model can make a repeated product feature viable once it meets a written quality bar. This page asks whether a 30% peak on a scientist-set science suite is a shipping research assistant. Both notes treat cost as part of the result. They do not share a workload.
The small-models note can stay with one person’s news-page experiment and listed token prices. This note stays with Dillmann’s scientist-set bar, the 70-task funnel, and the cost and token frontiers the suite already publishes. Collapsing those questions into one “cheaper is better” headline would drop the bar, the harness, and the task set.
How to read Terminal-Bench-Science
Read Dillmann’s announcement for the task funnel, the 70-task coverage, the named resolution rates, the cost and token frontiers, and the living-benchmark roadmap. Read the Hraness reading note of that announcement for a dated digest. Read the Hraness harness definition when a row’s agent name needs a noun. The next release, 0.2, has a pull-request deadline of October 5, 2026. Tasks are versioned so Harbor can re-run trials.
The useful sentence is narrower than a leaderboard headline. Scientists set the bar. The leading named configuration resolves 30% of the accepted workflows. That remaining miss rate is the result. Cost and token Pareto is how AI Charts already asks the next product question.
Limits of this reading
- Resolution rates, costs, and token totals belong to the named 0.1 release, models, harnesses, and three-trial protocol on the announcement. They can change in a later release.
- Terminal-Bench-Science 0.1 is not in the checked Artificial Analysis coding-agent snapshot. A stored Terminal-Bench v2.1 score is a different suite.
- Reported evaluation costs are totals across all 70 tasks. They are not a production invoice, a subscription price, or a per-query quote.
- The suite is a living benchmark. Later releases will add, retire, and recalibrate tasks as the frontier moves.
- Resolution varies by scientific domain. A suite score is not a domain score, and a domain lead is not a general research-assistant claim.
- The claim that 30% is not a product win is AI Charts analysis of those reported rates. Cite Dillmann for the measurements.
