Explore benchmarks
Choose a task, then compare models on the same test.
29 interactive charts · 62 benchmarks and guides
Coding / 4.0.0
Terminal-Bench 4
Which agent can finish difficult terminal work?
- Evidence
- Owner-published submissions
- Version
- 4.0.0
- Retrieved
- Configurations
- 13
Software, systems, CAD, scientific computing, and formal proof tasks executed in a terminal. Task success over 66 tasks and five trials per configuration, with source-reported 95% confidence intervals.
Showing 8 of 13 configurations.
Select a row to inspect the configuration or compare up to three. Whiskers: Source-reported 95% confidence interval.
What this chart shows
Fable 5.1 leads at 57.88%, ahead of Opus 5 at 51.82%. Their reported uncertainty ranges overlap, so this source does not separate the top two. Scores across the 13 charted systems run from 11.21% to 57.88%.
These sentences cover the whole chart and ignore the filters above. Each system counts once, at its best-scoring setting.
About these results. 66 tasks × 5 trials per configuration. Model, agent, and effort all affect results. The owner reports 95% confidence intervals; its interval method and price basis are unspecified.
What this benchmark measures and how to read it
Software, systems, CAD, scientific computing, and formal proof tasks executed in a terminal.
- Measure
- Task success over 66 tasks and five trials per configuration, with source-reported 95% confidence intervals.
- Compare fairly
- AI Charts’ primary terminal benchmark. Compare exact version 4.0.0 with the named model, agent version, and effort.
- An agent and model are evaluated together; this is not a model-only ranking.
- Different harnesses and effort settings remain separate configurations.
- Cost covers the full 330-trial evaluation. The source does not specify its pricing or confidence-interval method.