Efficient frontier
Models on the Pareto frontier by API cost per task and DeepSWE.
Performance, cost, and time for the full model-and-agent setup.
Loading chart…
Artificial Analysis · 68 configurations · 31 models · 11 providersSnapshot
DeepSWE — Long-horizon software engineering tasks scored with automated code verification.
Hover or select a point for exact valuesTap points
Option-space summary
Current read. GPT-5.6 Sol (max) reaches the highest sampled frontier score at 68.7; 11 provider ranges are available, led by OpenAI at 68.7.
Models on the Pareto frontier by API cost per task and DeepSWE.
Minimum, median, and maximum DeepSWE by provider.
Each point is a measured model-and-agent configuration from Artificial Analysis. Cost, time, and total tokens describe the full task run, not just the model’s output. Only configurations reporting both selected metrics appear in the chart.
This source reports Terminal-Bench v2.1. Its scores stay separate from Terminal-Bench 4, the current coding standard. A snapshot date records retrieval, not when each model was evaluated.
Daily snapshot diff · checked
5 models added
Fable 5.1 (with fallback) (Claude Code) · Muse Spark 1.3 (Muse Code) · GPT-6 Astra (Codex) · Gemini 3.8 Flash (Opencode) · Gemini 3.8 Flash (Antigravity SDK)
Claude Code · Anthropic · max
Muse Code · Meta · max · 2 settings
Codex · OpenAI · max
Opencode · Google · high
Antigravity SDK · Google · high
Claude Code · Anthropic · xhigh · 4 settings
Claude Code · Anthropic · none
Codex · OpenAI · high
Codex · OpenAI · xhigh
Claude Code · Alibaba Cloud · default
Devin CLI · Cognition · default
Codex · OpenAI · medium
Codex · DeepSeek · max