In the aicharts coding-agent snapshot retrieved Oct 5, 2026, 2:10 PM UTC, Codex · GPT-6.1 Sol (xhigh) scores 62.9 on AA Index at a mean API cost of $1.04 per task, fourth of the 31 configurations that carry an index. The max setting scores 60.1 at $1.55, so the highest Codex row is not max. OpenAI released GPT-6.1 Sol with the claim “Near-Astra intelligence for a fifth of the price,” and Artificial Analysis added the Codex rows to its coding-agents comparison afterwards. The snapshot’s update log records the configuration on October 5, 2026.
One model inside Codex, scored on three benchmarks
The coding-agent chart is a daily snapshot of the public Artificial Analysis coding-agents comparison. Each row is one configuration: a model, the agent harness that ran it, and an effort setting. Artificial Analysis runs the harness on tasks from three coding benchmarks, DeepSWE v1.1, Terminal-Bench 4, and SWE-Atlas-QnA, and AA Index is the mean of the three scores. The cost is the mean API bill for one task in that harness at list prices, including every tool call and every repeated read of the repository the harness sends.
GPT-6.1 Sol’s placed row is the model inside Codex at the xhigh setting. One task in that configuration used 3.2 million tokens and took 16 minutes of harness time on average. The same model at another setting is another row; the snapshot stores five configurations running GPT-6.1 Sol. OpenAI’s model page lists effort levels as low, medium (default), high, xhigh, and max, with a 1,050,000-token context window.
Fourth of 31 configurations
The 62.9 places Codex · GPT-6.1 Sol (xhigh) fourth of the 31 configurations that carry an AA Index in the snapshot retrieved Oct 5, 2026, 2:10 PM UTC. Three configurations score higher: Claude Code · Sonnet 5.5 (max) at 68.4, Claude Code · Opus 5.5 (max) at 66.0, and Antigravity CLI · Gemini 4 Argon (default) at 63.8. The leader, Claude Code · Sonnet 5.5 (max), is 5.5 points above it at $14.19 per task. Its $1.04 per task is the 22nd highest cost of the 31 configurations that carry a cost.
Claude Code · Sonnet 5.5 (xhigh) matches that printed score at 62.9 for 3.2x as much as the GPT-6.1 Sol row’s cost. The table lists the four highest-scoring configurations after it, with each cost as a share of its $1.04.
| Configuration | Setting | AA Index | Cost per task | Points below GPT-6.1 Sol | Share of GPT-6.1 Sol’s cost |
|---|---|---|---|---|---|
| Claude Code · Sonnet 5.5 | xhigh | 62.9 | $3.33 | Tie | 321% |
| Claude Code · Fable 5.1 (with fallback) | max | 62.2 | $12.39 | 0.7 points | 1191% |
| Devin Fusion CLI · Claude Fable 5.1 XHigh + SWE-2 Medium | default | 61.7 | $7.90 | 1.2 points | 760% |
| Codex · GPT-6 Astra | max | 61.6 | $7.47 | 1.3 points | 718% |
Five Codex settings, and max is not the top of them
The snapshot stores five Codex settings for GPT-6.1 Sol: low, medium, high, xhigh, and max. The placed row is xhigh, the highest AA Index among those settings. The max setting scores 60.1 at $1.55 per task, so it is neither the highest-scoring Codex row nor the cheapest.
| Setting | AA Index | Cost per task | On the cost frontier |
|---|---|---|---|
| low | 57.2 | $0.499 | Yes |
| medium | 61.4 | $0.705 | Yes |
| high | 60.1 | $0.889 | No |
| xhigh | 62.9 | $1.04 | Yes |
| max | 60.1 | $1.55 | No |
Low, medium, and xhigh sit on the cost frontier. high and max do not: another configuration costs no more and scores at least as high.
What stepping down the cost frontier gives up
The chart’s cost frontier contains configurations for which no other row costs no more and scores at least as high, with a strict improvement in either measure. Codex · GPT-6.1 Sol (xhigh) is on it. The question the frontier answers is what a reader gives up by stepping down from a higher score to a cheaper row that nothing dominates.
The two highest Claude Code rows cost much more for a higher score: Claude Code · Sonnet 5.5 (max) at 68.4 for $14.19 per task (13.6x the GPT-6.1 Sol cost), and Claude Code · Opus 5.5 (max) at 66.0 for $13.04 (12.5x). Those gaps are why a reader who cares about score per dollar stays with the Codex row rather than treating the chart as a single ranking.
| Configuration | AA Index | Cost per task |
|---|---|---|
| Codex · GPT-6.1 Sol (xhigh) | 62.9 | $1.04 |
| Claude Code · Sonnet 5.5 (max) | 68.4 | $14.19 |
| Claude Code · Opus 5.5 (max) | 66.0 | $13.04 |
The first step down is Codex · GPT-6.1 Sol (medium): 1.5 points lower for 68% of the $1.04. The first vertex at half the cost or less is Codex · GPT-6.1 Sol (low), which gives up 5.7 points for 48% of the cost. The frontier ends at Codex · DeepSeek V4 Flash 0731 (max): 24.2 points lower for 8.2% of the cost.
| Configuration | Setting | AA Index | Cost per task | Points below GPT-6.1 Sol | Share of GPT-6.1 Sol’s cost |
|---|---|---|---|---|---|
| Codex · GPT-6.1 Sol | medium | 61.4 | $0.705 | 1.5 points | 68% |
| Codex · GPT-6.1 Sol | low | 57.2 | $0.499 | 5.7 points | 48% |
| Codex · GPT-5.6 Luna | max | 43.2 | $0.438 | 19.7 points | 42% |
| Codex · DeepSeek V4 Pro 0813 | max | 43.1 | $0.238 | 19.9 points | 23% |
| Codex · GPT-6 Luna | max | 41.1 | $0.176 | 21.8 points | 17% |
| Codex · DeepSeek V4 Flash 0731 | max | 38.7 | $0.085 | 24.2 points | 8.2% |
Where the index points come from
AA Index is the mean of three component benchmarks, and the GPT-6.1 Sol row does not lead all three. On DeepSWE v1.1 it scores 73.2, second of 31; on Terminal-Bench 4 it scores 54.5, eighth of 31; and on SWE-Atlas-QnA it scores 61.0, 12th of 31 configurations that carry each score.
It leads none of the components outright; its composite position comes from placing high on all of them at once. Antigravity CLI · Gemini 4 Argon (default) scores 5.6 points higher on DeepSWE v1.1, Claude Code · Sonnet 5.5 (max) scores 11.6 points higher on Terminal-Bench 4, and Claude Code · Sonnet 5.5 (max) scores 5.9 points higher on SWE-Atlas-QnA.
| Component | GPT-6.1 Sol score | Rank | Best other configuration | Gap |
|---|---|---|---|---|
| DeepSWE v1.1 | 73.2 | 2 of 31 | Antigravity CLI · Gemini 4 Argon (default) at 78.8 | −5.6 points |
| Terminal-Bench 4 | 54.5 | 8 of 31 | Claude Code · Sonnet 5.5 (max) at 66.2 | −11.6 points |
| SWE-Atlas-QnA | 61.0 | 12 of 31 | Claude Code · Sonnet 5.5 (max) at 66.9 | −5.9 points |
Four configurations with a lower AA Index score higher on Terminal-Bench 4: Claude Code · Sonnet 5.5 (xhigh) at 58.1, Claude Code · Fable 5.1 (with fallback) (max) at 57.6, Devin Fusion CLI · Claude Fable 5.1 XHigh + SWE-2 Medium (default) at 56.1, and Codex · GPT-6 Astra (max) at 55.6. A reader who cares only about terminal work would order these rows differently from the composite.
GPT-6 Sol to GPT-6.1 Sol in Codex
The snapshot still stores the previous Sol generation in Codex: Codex · GPT-6 Sol (max) at 56.7 for $2.99 per task. The highest-scoring GPT-6.1 Sol row is Codex · GPT-6.1 Sol (xhigh), a different setting from Codex · GPT-6 Sol (max), so the generation step below is not a like-for-like rerun. GPT-6.1 Sol adds 6.2 points at 0.3x the mean cost per task. At the shared max setting, Codex · GPT-6.1 Sol (max) scores 60.1 at $1.55 per task.
| Measure | GPT-6 Sol | GPT-6.1 Sol | Change |
|---|---|---|---|
| AA Index | 56.7 | 62.9 | +6.2 points |
| DeepSWE v1.1 | 69.0 | 73.2 | +4.1 points |
| Terminal-Bench 4 | 43.4 | 54.5 | +11.1 points |
| SWE-Atlas-QnA | 57.5 | 61.0 | +3.5 points |
| Mean API cost per task | $2.99 | $1.04 | 0.3x |
The earlier GPT-6 Sol note is pinned to the September 25, 2026 coding-agent snapshot and the September 23, 2026 Intelligence Index snapshot it cited. That note’s Index figures describe a roster that no longer lists GPT-6 Sol. This page reads the live coding-agent snapshot and uses the live Index only for the cost callout below.
The Intelligence Index row is a different measurement
The Artificial Analysis Intelligence Index runs the model through its API under one harness that is the same for every model, across 10 evaluations at version 4.3.2, and its cost per task is the average bill for one of those evaluation tasks. In the snapshot retrieved Oct 6, 2026, 11:15 AM UTC, GPT-6.1 Sol (Max) scores 51.8 at $0.724 per task. That row is a different unit from the Codex AA Index: this note does not rank it with coding-agent scores or walk the Index frontier.
| Chart | Configuration | Score | Cost per task | Snapshot retrieved |
|---|---|---|---|---|
| Coding agents (AA Index) | Codex · GPT-6.1 Sol (xhigh) | 62.9 | $1.04 | Oct 5, 2026, 2:10 PM UTC |
| Intelligence Index | GPT-6.1 Sol (Max) | 51.8 | $0.724 | Oct 6, 2026, 11:15 AM UTC |
The 62.9 is a mean of three coding benchmarks run inside Codex, and the 51.8 is a weighted average of 10 evaluations run through the API. The $1.04 includes every tool call and every repeated read of the repository that the harness sends, 3.2 million tokens per task in this snapshot; the $0.724 is the average bill for one evaluation task under Artificial Analysis’s standardized harness, with 38,128 output tokens per task. Artificial Analysis’s model page, captured October 5, 2026 UTC, lists the same list prices behind both figures, $2.00 per million input tokens and $10.00 per million output tokens with a 95% cache discount, a rounded index of 52, and 67M output tokens to run the whole index. The page’s FAQ dates the release to September 29, 2026.
OpenAI’s own figures
OpenAI’s launch page opens with “Near-Astra intelligence for a fifth of the price” and reports its own DeepSWE v1.1 run: GPT-6.1 Sol “matches GPT‑6 Astra at roughly one-fifth of the cost, while eclipsing GPT‑6 Sol’s best score by 6.4 percentage points at a lower reasoning effort and cost.” It also prints vendor-run results on GDP.pdf, AutomationBench, OSWorld 2.0, and Terminal-Bench Science 0.1, none of which appears on an aicharts chart. Those figures come from OpenAI’s evaluation setup, not from Codex under Artificial Analysis’s protocol. Read them as the vendor’s description of its model and read the chart for an independent measurement of one named configuration.
The same page prices the API at $2 per million input tokens, $10 per million output tokens, and $0.10 per million cached input tokens, and makes the model available to Plus, Pro, Business, Enterprise, and Edu users in ChatGPT Work and Codex. It also states that “GPT‑6.1 Sol is not yet available in Chat.” Developers reach it through the API as gpt-6.1-sol. Artificial Analysis’s model page FAQ dates that release to September 29, 2026 and describes the model as “amongst the leading models in intelligence and reasonably priced when comparing to other models of similar price.”
Limits
- The chart scores, task costs, token counts, and durations are Artificial Analysis measurements of the named configuration on the retrieval date, under DeepSWE v1.1, Terminal-Bench 4, and SWE-Atlas-QnA for the coding-agent chart and Intelligence Index version 4.3.2 for the capability chart. None of them establishes a result on other repositories, tasks, or harnesses.
- The rank, cost rank, frontier steps, component gaps, and generation multiples are aicharts derivations from the snapshots named in each caption. A configuration added, removed, or rescored by Artificial Analysis moves them, and the coding-agent snapshot advances daily.
- GPT-6.1 Sol in Cursor, Devin, or any harness other than Codex is a configuration this snapshot does not store, so this note says nothing about it.
- The GPT-6 Sol comparison may hold the harness fixed while the settings differ. The snapshot records outcomes, not run dates or benchmark versions at run time; Artificial Analysis may have measured the two generations days apart.
- The Intelligence Index row is a second measurement of the same model, not a second coding-agent score. GPT-6 Sol is absent from the live Index roster; the earlier Sol note keeps the September 23, 2026 Index snapshot it cited.
- The prices, the availability statement, and the vendor-run benchmark figures belong to OpenAI. aicharts did not run GPT-6.1 Sol.
