Explore benchmarks
Choose a task, then compare models on the same test.
24 interactive charts · 57 benchmarks and guides
Coding / 4.0.0
Terminal-Bench 4
Which agent can finish difficult terminal work?
13 configurations in source cohortOwner-published submissionsRetrieved Sep 3, 2026
Showing 8 of 13 configurations.
Model, agent version, and effortTask success · higher is better
Select a row to inspect the configuration or compare up to three. Whiskers: Source-reported 95% confidence interval.
About these results. 66 tasks × 5 trials per configuration. Model, agent, and effort all affect results. The owner reports 95% confidence intervals; its interval method and price basis are unspecified.
Harbor Framework ↗Retrieved
What this benchmark measures and how to read it
Software, systems, CAD, scientific computing, and formal proof tasks executed in a terminal.
- Measure
- Task success over 66 tasks and five trials per configuration, with source-reported 95% confidence intervals.
- Compare fairly
- AI Charts’ primary terminal benchmark. Compare exact version 4.0.0 with the named model, agent version, and effort.
- An agent and model are evaluated together; this is not a model-only ranking.
- Different harnesses and effort settings remain separate configurations.
- Cost covers the full 330-trial evaluation. The source does not specify its pricing or confidence-interval method.