Skip to content

Explore benchmarks

Choose a task, then compare models on the same test.

29 interactive charts · 62 benchmarks and guides

Task
Benchmark
Browse library
62 benchmarks

Charts include results you can explore here. Guides link to published evaluations; emerging benchmarks do not yet have comparable scores here.

Coding / 4.0.0

Terminal-Bench 4

Which agent can finish difficult terminal work?

Evidence
Owner-published submissions
Version
4.0.0
Retrieved
Configurations
13

Software, systems, CAD, scientific computing, and formal proof tasks executed in a terminal. Task success over 66 tasks and five trials per configuration, with source-reported 95% confidence intervals.

Filters
Provider

Showing 8 of 13 configurations.

Model, agent version, and effortTask success · higher is better

Select a row to inspect the configuration or compare up to three. Whiskers: Source-reported 95% confidence interval.

What this chart shows

Fable 5.1 leads at 57.88%, ahead of Opus 5 at 51.82%. Their reported uncertainty ranges overlap, so this source does not separate the top two. Scores across the 13 charted systems run from 11.21% to 57.88%.

These sentences cover the whole chart and ignore the filters above. Each system counts once, at its best-scoring setting.

About these results. 66 tasks × 5 trials per configuration. Model, agent, and effort all affect results. The owner reports 95% confidence intervals; its interval method and price basis are unspecified.

Harbor Framework ↗Retrieved
What this benchmark measures and how to read it

Software, systems, CAD, scientific computing, and formal proof tasks executed in a terminal.

Measure
Task success over 66 tasks and five trials per configuration, with source-reported 95% confidence intervals.
Compare fairly
AI Charts’ primary terminal benchmark. Compare exact version 4.0.0 with the named model, agent version, and effort.
  • An agent and model are evaluated together; this is not a model-only ranking.
  • Different harnesses and effort settings remain separate configurations.
  • Cost covers the full 330-trial evaluation. The source does not specify its pricing or confidence-interval method.
Read the methodology ↗