Skip to notes
aicharts
Theme
Appearance

GPT-6.1 Sol’s best Codex row is 62.9 at $1.04 a task, not max

Codex · GPT-6.1 Sol (xhigh) scores 62.9 on AA Index at $1.04 per task, fourth of 31 configurations. Max scores 60.1 at $1.55.

Drafted with AI and reviewed by Codex.

Five ivory discs rest on a short charcoal rail; the fourth is the largest, and the fifth is smaller than the third.
Codex · GPT-6.1 Sol’s highest AA Index is not its max setting, and not every Codex setting sits on the coding-agent cost frontier. Made with SlopCamera

In the aicharts coding-agent snapshot retrieved Oct 5, 2026, 2:10 PM UTC, Codex · GPT-6.1 Sol (xhigh) scores 62.9 on AA Index at a mean API cost of $1.04 per task, fourth of the 31 configurations that carry an index. The max setting scores 60.1 at $1.55, so the highest Codex row is not max. OpenAI released GPT-6.1 Sol with the claim “Near-Astra intelligence for a fifth of the price,” and Artificial Analysis added the Codex rows to its coding-agents comparison afterwards. The snapshot’s update log records the configuration on October 5, 2026.

One model inside Codex, scored on three benchmarks

The coding-agent chart is a daily snapshot of the public Artificial Analysis coding-agents comparison. Each row is one configuration: a model, the agent harness that ran it, and an effort setting. Artificial Analysis runs the harness on tasks from three coding benchmarks, DeepSWE v1.1, Terminal-Bench 4, and SWE-Atlas-QnA, and AA Index is the mean of the three scores. The cost is the mean API bill for one task in that harness at list prices, including every tool call and every repeated read of the repository the harness sends.

GPT-6.1 Sol’s placed row is the model inside Codex at the xhigh setting. One task in that configuration used 3.2 million tokens and took 16 minutes of harness time on average. The same model at another setting is another row; the snapshot stores five configurations running GPT-6.1 Sol. OpenAI’s model page lists effort levels as low, medium (default), high, xhigh, and max, with a 1,050,000-token context window.

Fourth of 31 configurations

The 62.9 places Codex · GPT-6.1 Sol (xhigh) fourth of the 31 configurations that carry an AA Index in the snapshot retrieved Oct 5, 2026, 2:10 PM UTC. Three configurations score higher: Claude Code · Sonnet 5.5 (max) at 68.4, Claude Code · Opus 5.5 (max) at 66.0, and Antigravity CLI · Gemini 4 Argon (default) at 63.8. The leader, Claude Code · Sonnet 5.5 (max), is 5.5 points above it at $14.19 per task. Its $1.04 per task is the 22nd highest cost of the 31 configurations that carry a cost.

Claude Code · Sonnet 5.5 (xhigh) matches that printed score at 62.9 for 3.2x as much as the GPT-6.1 Sol row’s cost. The table lists the four highest-scoring configurations after it, with each cost as a share of its $1.04.

The four highest-scoring configurations after Codex · GPT-6.1 Sol (xhigh) in the snapshot retrieved Oct 5, 2026, 2:10 PM UTC
ConfigurationSettingAA IndexCost per taskPoints below GPT-6.1 SolShare of GPT-6.1 Sol’s cost
Claude Code · Sonnet 5.5xhigh62.9$3.33Tie321%
Claude Code · Fable 5.1 (with fallback)max62.2$12.390.7 points1191%
Devin Fusion CLI · Claude Fable 5.1 XHigh + SWE-2 Mediumdefault61.7$7.901.2 points760%
Codex · GPT-6 Astramax61.6$7.471.3 points718%

Five Codex settings, and max is not the top of them

The snapshot stores five Codex settings for GPT-6.1 Sol: low, medium, high, xhigh, and max. The placed row is xhigh, the highest AA Index among those settings. The max setting scores 60.1 at $1.55 per task, so it is neither the highest-scoring Codex row nor the cheapest.

Codex · GPT-6.1 Sol settings in the snapshot retrieved Oct 5, 2026, 2:10 PM UTC, cheapest setting first
SettingAA IndexCost per taskOn the cost frontier
low57.2$0.499Yes
medium61.4$0.705Yes
high60.1$0.889No
xhigh62.9$1.04Yes
max60.1$1.55No

Low, medium, and xhigh sit on the cost frontier. high and max do not: another configuration costs no more and scores at least as high.

What stepping down the cost frontier gives up

The chart’s cost frontier contains configurations for which no other row costs no more and scores at least as high, with a strict improvement in either measure. Codex · GPT-6.1 Sol (xhigh) is on it. The question the frontier answers is what a reader gives up by stepping down from a higher score to a cheaper row that nothing dominates.

The two highest Claude Code rows cost much more for a higher score: Claude Code · Sonnet 5.5 (max) at 68.4 for $14.19 per task (13.6x the GPT-6.1 Sol cost), and Claude Code · Opus 5.5 (max) at 66.0 for $13.04 (12.5x). Those gaps are why a reader who cares about score per dollar stays with the Codex row rather than treating the chart as a single ranking.

Codex · GPT-6.1 Sol (xhigh) against the highest Claude Code rows in the snapshot retrieved Oct 5, 2026, 2:10 PM UTC
ConfigurationAA IndexCost per task
Codex · GPT-6.1 Sol (xhigh)62.9$1.04
Claude Code · Sonnet 5.5 (max)68.4$14.19
Claude Code · Opus 5.5 (max)66.0$13.04

The first step down is Codex · GPT-6.1 Sol (medium): 1.5 points lower for 68% of the $1.04. The first vertex at half the cost or less is Codex · GPT-6.1 Sol (low), which gives up 5.7 points for 48% of the cost. The frontier ends at Codex · DeepSeek V4 Flash 0731 (max): 24.2 points lower for 8.2% of the cost.

Cost-frontier configurations below Codex · GPT-6.1 Sol (xhigh) in the snapshot retrieved Oct 5, 2026, 2:10 PM UTC, highest AA Index first
ConfigurationSettingAA IndexCost per taskPoints below GPT-6.1 SolShare of GPT-6.1 Sol’s cost
Codex · GPT-6.1 Solmedium61.4$0.7051.5 points68%
Codex · GPT-6.1 Sollow57.2$0.4995.7 points48%
Codex · GPT-5.6 Lunamax43.2$0.43819.7 points42%
Codex · DeepSeek V4 Pro 0813max43.1$0.23819.9 points23%
Codex · GPT-6 Lunamax41.1$0.17621.8 points17%
Codex · DeepSeek V4 Flash 0731max38.7$0.08524.2 points8.2%

Where the index points come from

AA Index is the mean of three component benchmarks, and the GPT-6.1 Sol row does not lead all three. On DeepSWE v1.1 it scores 73.2, second of 31; on Terminal-Bench 4 it scores 54.5, eighth of 31; and on SWE-Atlas-QnA it scores 61.0, 12th of 31 configurations that carry each score.

It leads none of the components outright; its composite position comes from placing high on all of them at once. Antigravity CLI · Gemini 4 Argon (default) scores 5.6 points higher on DeepSWE v1.1, Claude Code · Sonnet 5.5 (max) scores 11.6 points higher on Terminal-Bench 4, and Claude Code · Sonnet 5.5 (max) scores 5.9 points higher on SWE-Atlas-QnA.

Codex · GPT-6.1 Sol (xhigh) on each AA Index component in the snapshot retrieved Oct 5, 2026, 2:10 PM UTC, against the best other configuration that carries the component
ComponentGPT-6.1 Sol scoreRankBest other configurationGap
DeepSWE v1.173.22 of 31Antigravity CLI · Gemini 4 Argon (default) at 78.8−5.6 points
Terminal-Bench 454.58 of 31Claude Code · Sonnet 5.5 (max) at 66.2−11.6 points
SWE-Atlas-QnA61.012 of 31Claude Code · Sonnet 5.5 (max) at 66.9−5.9 points

Four configurations with a lower AA Index score higher on Terminal-Bench 4: Claude Code · Sonnet 5.5 (xhigh) at 58.1, Claude Code · Fable 5.1 (with fallback) (max) at 57.6, Devin Fusion CLI · Claude Fable 5.1 XHigh + SWE-2 Medium (default) at 56.1, and Codex · GPT-6 Astra (max) at 55.6. A reader who cares only about terminal work would order these rows differently from the composite.

GPT-6 Sol to GPT-6.1 Sol in Codex

The snapshot still stores the previous Sol generation in Codex: Codex · GPT-6 Sol (max) at 56.7 for $2.99 per task. The highest-scoring GPT-6.1 Sol row is Codex · GPT-6.1 Sol (xhigh), a different setting from Codex · GPT-6 Sol (max), so the generation step below is not a like-for-like rerun. GPT-6.1 Sol adds 6.2 points at 0.3x the mean cost per task. At the shared max setting, Codex · GPT-6.1 Sol (max) scores 60.1 at $1.55 per task.

Codex · GPT-6 Sol and Codex · GPT-6.1 Sol (xhigh) in the snapshot retrieved Oct 5, 2026, 2:10 PM UTC
MeasureGPT-6 SolGPT-6.1 SolChange
AA Index56.762.9+6.2 points
DeepSWE v1.169.073.2+4.1 points
Terminal-Bench 443.454.5+11.1 points
SWE-Atlas-QnA57.561.0+3.5 points
Mean API cost per task$2.99$1.040.3x

The earlier GPT-6 Sol note is pinned to the September 25, 2026 coding-agent snapshot and the September 23, 2026 Intelligence Index snapshot it cited. That note’s Index figures describe a roster that no longer lists GPT-6 Sol. This page reads the live coding-agent snapshot and uses the live Index only for the cost callout below.

The Intelligence Index row is a different measurement

The Artificial Analysis Intelligence Index runs the model through its API under one harness that is the same for every model, across 10 evaluations at version 4.3.2, and its cost per task is the average bill for one of those evaluation tasks. In the snapshot retrieved Oct 6, 2026, 11:15 AM UTC, GPT-6.1 Sol (Max) scores 51.8 at $0.724 per task. That row is a different unit from the Codex AA Index: this note does not rank it with coding-agent scores or walk the Index frontier.

GPT-6.1 Sol in the two aicharts snapshots, each on its own task set with its own cost definition
ChartConfigurationScoreCost per taskSnapshot retrieved
Coding agents (AA Index)Codex · GPT-6.1 Sol (xhigh)62.9$1.04Oct 5, 2026, 2:10 PM UTC
Intelligence IndexGPT-6.1 Sol (Max)51.8$0.724Oct 6, 2026, 11:15 AM UTC

The 62.9 is a mean of three coding benchmarks run inside Codex, and the 51.8 is a weighted average of 10 evaluations run through the API. The $1.04 includes every tool call and every repeated read of the repository that the harness sends, 3.2 million tokens per task in this snapshot; the $0.724 is the average bill for one evaluation task under Artificial Analysis’s standardized harness, with 38,128 output tokens per task. Artificial Analysis’s model page, captured October 5, 2026 UTC, lists the same list prices behind both figures, $2.00 per million input tokens and $10.00 per million output tokens with a 95% cache discount, a rounded index of 52, and 67M output tokens to run the whole index. The page’s FAQ dates the release to September 29, 2026.

OpenAI’s own figures

OpenAI’s launch page opens with “Near-Astra intelligence for a fifth of the price” and reports its own DeepSWE v1.1 run: GPT-6.1 Sol “matches GPT‑6 Astra at roughly one-fifth of the cost, while eclipsing GPT‑6 Sol’s best score by 6.4 percentage points at a lower reasoning effort and cost.” It also prints vendor-run results on GDP.pdf, AutomationBench, OSWorld 2.0, and Terminal-Bench Science 0.1, none of which appears on an aicharts chart. Those figures come from OpenAI’s evaluation setup, not from Codex under Artificial Analysis’s protocol. Read them as the vendor’s description of its model and read the chart for an independent measurement of one named configuration.

The same page prices the API at $2 per million input tokens, $10 per million output tokens, and $0.10 per million cached input tokens, and makes the model available to Plus, Pro, Business, Enterprise, and Edu users in ChatGPT Work and Codex. It also states that “GPT‑6.1 Sol is not yet available in Chat.” Developers reach it through the API as gpt-6.1-sol. Artificial Analysis’s model page FAQ dates that release to September 29, 2026 and describes the model as “amongst the leading models in intelligence and reasonably priced when comparing to other models of similar price.”

Limits

  • The chart scores, task costs, token counts, and durations are Artificial Analysis measurements of the named configuration on the retrieval date, under DeepSWE v1.1, Terminal-Bench 4, and SWE-Atlas-QnA for the coding-agent chart and Intelligence Index version 4.3.2 for the capability chart. None of them establishes a result on other repositories, tasks, or harnesses.
  • The rank, cost rank, frontier steps, component gaps, and generation multiples are aicharts derivations from the snapshots named in each caption. A configuration added, removed, or rescored by Artificial Analysis moves them, and the coding-agent snapshot advances daily.
  • GPT-6.1 Sol in Cursor, Devin, or any harness other than Codex is a configuration this snapshot does not store, so this note says nothing about it.
  • The GPT-6 Sol comparison may hold the harness fixed while the settings differ. The snapshot records outcomes, not run dates or benchmark versions at run time; Artificial Analysis may have measured the two generations days apart.
  • The Intelligence Index row is a second measurement of the same model, not a second coding-agent score. GPT-6 Sol is absent from the live Index roster; the earlier Sol note keeps the September 23, 2026 Index snapshot it cited.
  • The prices, the availability statement, and the vendor-run benchmark figures belong to OpenAI. aicharts did not run GPT-6.1 Sol.

Sources

  1. Coding AgentsArtificial Analysis, 2026. The public coding-agents comparison is the source of the aicharts coding-agent snapshot. Model names, agent harnesses, settings, AA Index scores, and mean API costs are Artificial Analysis measurements.
  2. Introducing GPT-6.1 SolOpenAI, 2026. Cited for the launch headline, the $2, $10, and $0.10 per million token prices, ChatGPT Work and Codex availability, the statement that the model is not yet in Chat, the API name, and the vendor-run DeepSWE v1.1 claim against GPT-6 Astra and GPT-6 Sol. aicharts does not chart those vendor figures.
  3. GPT-6.1 Sol ModelOpenAI, 2026. Cited for the effort levels low, medium (default), high, xhigh, and max, the 1,050,000-token context window, and the matching list prices.
  4. LLM LeaderboardArtificial Analysis, 2026. The public models leaderboard is the source of the aicharts Intelligence Index snapshot. Scores, per-task costs, and output tokens are Artificial Analysis measurements under Intelligence Index v4.3.2.
  5. GPT-6.1 Sol (max) - Intelligence, Performance & Price AnalysisArtificial Analysis, 2026. Cited, from the page captured October 5, 2026 UTC, for the rounded 52 index score, the proprietary label, the September 29, 2026 release date, the $2.00 and $10.00 per million token prices with a 95% cache discount, the $0.72 cost per index task, the 67M output tokens across the index, and the 1M token context window.

Figures come from the cited primary sources and the aicharts datasets. aicharts did not rerun the reported benchmarks.