Skip to content

Explore benchmarks

Choose a task, then compare models on the same test.

24 interactive charts · 57 benchmarks and guides

Browse library
57 benchmarks

Charts include results you can explore here. Guides link to published evaluations; emerging benchmarks do not yet have comparable scores here.

Coding / 4.0.0

Terminal-Bench 4

Which agent can finish difficult terminal work?

13 configurations in source cohortOwner-published submissionsRetrieved Sep 3, 2026
Filters

Showing 8 of 13 configurations.

Model, agent version, and effortTask success · higher is better

Select a row to inspect the configuration or compare up to three. Whiskers: Source-reported 95% confidence interval.

About these results. 66 tasks × 5 trials per configuration. Model, agent, and effort all affect results. The owner reports 95% confidence intervals; its interval method and price basis are unspecified.

Harbor FrameworkRetrieved
What this benchmark measures and how to read it

Software, systems, CAD, scientific computing, and formal proof tasks executed in a terminal.

Measure
Task success over 66 tasks and five trials per configuration, with source-reported 95% confidence intervals.
Compare fairly
AI Charts’ primary terminal benchmark. Compare exact version 4.0.0 with the named model, agent version, and effort.
  • An agent and model are evaluated together; this is not a model-only ranking.
  • Different harnesses and effort settings remain separate configurations.
  • Cost covers the full 330-trial evaluation. The source does not specify its pricing or confidence-interval method.
Read the methodology ↗