In the aicharts coding-agent snapshot retrieved Oct 5, 2026, 2:10 PM UTC, Antigravity CLI · Gemini 4 Argon (default) scores 63.8 on AA Index at a mean API cost of $5.84 per task, third of the 31 configurations that carry an index. Two configurations score higher, and the cheaper of them costs 2.2x as much per task. Google announced Gemini 4 Argon on September 30, 2026 as a model “rolling out to a set of trusted cyber defenders through our Fairwind Program,” not as a general release. The post describes access in terms of those defenders, trusted testers, Google’s own internal teams, and the U.S. government’s voluntary pre-release process, gives no date for wider access, and Artificial Analysis’s model page marks the model “Not publicly available” on its charts. The chart row this note reads is an independent measurement of a model most readers cannot yet buy. The snapshot’s update log records the row on October 5, 2026.
One model inside one harness, scored on three benchmarks
The coding-agent chart is a daily snapshot of the public Artificial Analysis coding-agents comparison. Each row is one configuration: a model, the agent harness that ran it, and an effort setting. Artificial Analysis runs the harness on tasks from three coding benchmarks, DeepSWE v1.1, Terminal-Bench 4, and SWE-Atlas-QnA, and AA Index is the mean of the three scores. The cost is the mean API bill for one task in that harness at list prices, including every tool call and every repeated read of the repository the harness sends.
Gemini 4 Argon’s row is the model inside Antigravity CLI, Google’s coding agent, at the default setting. One task in that configuration used 13.7 million tokens and took 35 minutes of harness time on average. The snapshot stores one configuration running Gemini 4 Argon, so there is no effort ladder to climb and no second harness to compare; the same model at another setting or in another harness would be another row. The other costed Google row on the snapshot is Antigravity SDK · Gemini 3.8 Flash (high) at 41.9 for $2.47: a different model in a different harness, so the snapshot offers no same-harness predecessor for Argon.
Third of 31 configurations
The 63.8 places Antigravity CLI · Gemini 4 Argon (default) third of the 31 configurations that carry an AA Index. Two configurations score higher: Claude Code · Sonnet 5.5 (max) at 68.4 and Claude Code · Opus 5.5 (max) at 66.0. The leader, Claude Code · Sonnet 5.5 (max), is 4.6 points above it at $14.19 per task. Its $5.84 per task is the ninth highest cost of the 31 configurations that carry a cost; six configurations that score lower cost more per task.
The nearest score below it is Codex · GPT-6.1 Sol (xhigh) at 62.9, 0.8 points lower for 18% of the Argon row’s cost. The table lists the four highest-scoring configurations after it, with each cost as a share of its $5.84.
| Configuration | Setting | AA Index | Cost per task | Points below Argon | Share of Argon’s cost |
|---|---|---|---|---|---|
| Codex · GPT-6.1 Sol | xhigh | 62.9 | $1.04 | 0.8 points | 18% |
| Claude Code · Sonnet 5.5 | xhigh | 62.9 | $3.33 | 0.9 points | 57% |
| Claude Code · Fable 5.1 (with fallback) | max | 62.2 | $12.39 | 1.5 points | 212% |
| Devin Fusion CLI · Claude Fable 5.1 XHigh + SWE-2 Medium | default | 61.7 | $7.90 | 2.1 points | 135% |
Every higher row costs at least twice as much
The chart’s horizontal axis is cost per task, so the question for a row that is not first is what the rows above it charge for their extra points. Every configuration that scores above Argon costs more per task. Claude Code · Sonnet 5.5 (max) charges $14.19 for 4.6 points more, so the $5.84 is 41% of its cost. Claude Code · Opus 5.5 (max) charges $13.04 for 2.2 points more, so the $5.84 is 45% of its cost. The cheapest way to buy a higher score on this snapshot is Claude Code · Opus 5.5 (max) at 2.2x the Argon row’s cost.
| Configuration | Setting | AA Index | Cost per task | Points above Argon | Argon’s cost as a share of theirs |
|---|---|---|---|---|---|
| Claude Code · Sonnet 5.5 | max | 68.4 | $14.19 | 4.6 points | 41% |
| Claude Code · Opus 5.5 | max | 66.0 | $13.04 | 2.2 points | 45% |
A reader who wants a score above 63.8 on this snapshot pays at least 2.2x the Argon row’s cost per task to get it. That is the row’s position in one sentence: not the highest score, and at most half the cost of any row that scores above it.
The frontier row beneath it gives up less than a point
The chart’s cost frontier contains configurations for which no other row costs no more and scores at least as high, with a strict improvement in either measure. Antigravity CLI · Gemini 4 Argon (default) is on it: nothing on the snapshot scores as high for the same money or less. The question the frontier answers is what a reader gives up by stepping down from this score to a cheaper row that nothing dominates.
The first step down is Codex · GPT-6.1 Sol (xhigh): 0.8 points lower for 18% of the $5.84. That is the step a cost-minded reader will notice: less than a point of index for less than a quarter of the cost. The Argon row’s claim on the frontier is the score itself, not a cost advantage over the row beneath it. The frontier ends at Codex · DeepSeek V4 Flash 0731 (max): 25.0 points lower for 1.5% of the cost.
| Configuration | Setting | AA Index | Cost per task | Points below Argon | Share of Argon’s cost |
|---|---|---|---|---|---|
| Codex · GPT-6.1 Sol | xhigh | 62.9 | $1.04 | 0.8 points | 18% |
| Codex · GPT-6.1 Sol | medium | 61.4 | $0.705 | 2.3 points | 12% |
| Codex · GPT-6.1 Sol | low | 57.2 | $0.499 | 6.5 points | 8.5% |
| Codex · GPT-5.6 Luna | max | 43.2 | $0.438 | 20.5 points | 7.5% |
| Codex · DeepSeek V4 Pro 0813 | max | 43.1 | $0.238 | 20.7 points | 4.1% |
| Codex · GPT-6 Luna | max | 41.1 | $0.176 | 22.7 points | 3.0% |
| Codex · DeepSeek V4 Flash 0731 | max | 38.7 | $0.085 | 25.0 points | 1.5% |
DeepSWE v1.1 carries the composite
AA Index is the mean of three component benchmarks, each ranked here among the configurations that carry it, and the Argon row leads one of them. On DeepSWE v1.1 it scores 78.8, first of 31; on Terminal-Bench 4 it scores 56.1, fifth of 31, tied with Devin Fusion CLI · Claude Fable 5.1 XHigh + SWE-2 Medium (default); and on SWE-Atlas-QnA it scores 56.5, 21st of 31, tied with Claude Code · Sonnet 5.5 (high).
It leads DeepSWE v1.1 by 5.6 points over Muse Code · Muse Spark 1.3 (xhigh) and Codex · GPT-6.1 Sol (xhigh), tied. Claude Code · Sonnet 5.5 (max) scores 10.1 points higher on Terminal-Bench 4 and 10.5 points higher on SWE-Atlas-QnA. The third place is a DeepSWE v1.1 result: the row is fifth of 31 on Terminal-Bench 4 and 21st of 31 on SWE-Atlas-QnA, and the composite rests on the component it leads.
| Component | Argon score | Rank | Best other configuration | Gap |
|---|---|---|---|---|
| DeepSWE v1.1 | 78.8 | 1 of 31 | Muse Code · Muse Spark 1.3 (xhigh) at 73.2 | +5.6 points |
| Terminal-Bench 4 | 56.1 | 5 of 31 | Claude Code · Sonnet 5.5 (max) at 66.2 | −10.1 points |
| SWE-Atlas-QnA | 56.5 | 21 of 31 | Claude Code · Sonnet 5.5 (max) at 66.9 | −10.5 points |
Two configurations with a lower AA Index score higher on Terminal-Bench 4: Claude Code · Sonnet 5.5 (xhigh) at 58.1 and Claude Code · Fable 5.1 (with fallback) (max) at 57.6. A reader who cares only about terminal work would order these rows differently from the composite.
The Index row is a different unit
The Artificial Analysis Intelligence Index runs the model through its API under one harness that is the same for every model, across 10 evaluations at version 4.3.2, and its cost per task is the average bill for one of those evaluation tasks. In the snapshot retrieved Oct 6, 2026, 11:15 AM UTC, Gemini 4 Argon (High) scores 52.6 at $1.99 per task, eighth of 103 configurations with a measured cost.
| Chart | Configuration | Score | Rank | Cost per task | Snapshot retrieved |
|---|---|---|---|---|---|
| Coding agents (AA Index) | Antigravity CLI · Gemini 4 Argon (default) | 63.8 | 3 of 31 | $5.84 | Oct 5, 2026, 2:10 PM UTC |
| Intelligence Index | Gemini 4 Argon (High) | 52.6 | 8 of 103 | $1.99 | Oct 6, 2026, 11:15 AM UTC |
The 63.8 is a mean of three coding benchmarks run inside Antigravity CLI, and the 52.6 is a weighted average of 10 evaluations run through the API. The $5.84 is the harness’s full bill for one task, 13.7 million tokens per task in this snapshot; the $1.99 is the average bill for one evaluation task under Artificial Analysis’s standardized harness, with 61,558 output tokens per task. The model places third on the coding-agent chart and eighth on the Index. The two ranks come from different task sets, different cohorts, and different cost definitions. Neither rank corrects the other. Artificial Analysis’s model page, captured October 6, 2026 UTC, lists the list prices behind both figures, $2.00 per million input tokens and $10.00 per million output tokens with a 95% cache discount, a rounded index of 53 (#8 of 225 models), and 110M output tokens to run the whole index. The page’s FAQ dates the release to September 30, 2026.
Who can run the model, and at what price
Google’s announcement, titled “Gemini 4 Argon: our next era of frontier intelligence,” describes a model “rolling out to a set of trusted cyber defenders through our Fairwind Program.” It states that “Safely releasing frontier capabilities at this level requires a phased approach” and that Google will keep gathering feedback from early testers “before making Argon available to developers, enterprises, and consumers as soon as possible,” with that wider release “starting with paid API customers and Google AI Ultra subscribers.” No date is given. The same page says that “For trusted defenders and our own internal teams at Google, we’ll be releasing Argon without cyber guardrails.” A reader of the chart should hold both facts at once: the score is an independent measurement, and the configuration it measures is one that Artificial Analysis could reach and most readers cannot.
The announcement sets an introductory price of $2 per million input tokens and $10 per million output tokens, with “cached input tokens priced at 95% off input token price,” and its footnote states that after the introductory period “the price of $4 per 1M input tokens and $20 per 1M output tokens will apply.” The chart’s cost column is computed at list prices, so a list-price change moves the row horizontally while the score stays where it is. The same page raises the output token limit to “an industry-leading 1M tokens, up from the previous 64K tokens.” Artificial Analysis’s model page describes the model as “amongst the leading models in intelligence and reasonably priced when comparing to other models of similar price” with a 1M-token context window. The same page marks the model “Not publicly available” on its charts while its FAQ says the model is “available via API through 1 provider.” The two lines describe one situation: an API exists, Artificial Analysis measured the model through it, and Google’s post says who may use it.
Google’s own figures
Google reports its own DeepSWE v1.1 run: “It sets a new state of the art on DeepSWE v1.1 (77.9%).” Google does not say which harness, task sample, or date produced that figure, so it and the chart’s score are not one measurement taken twice. The chart’s DeepSWE v1.1 score for the Antigravity CLI row is 78.8, measured by Artificial Analysis inside Antigravity CLI on the retrieval date. The announcement also prints vendor-run results on the Vals Index, AutomationBench, CWE-bench v1, and LVBench, none of which appears on the coding-agent chart. AutomationBench-AA is Artificial Analysis’s own variant of one of them and counts toward the Index row above, not toward the coding-agent row. Read Google’s figures as the vendor’s description of its model and read the chart for an independent measurement of one named configuration.
Limits
- The chart scores, task costs, token counts, and durations are Artificial Analysis measurements of the named configuration on the retrieval date, under DeepSWE v1.1, Terminal-Bench 4, and SWE-Atlas-QnA for the coding-agent chart and Intelligence Index version 4.3.2 for the capability chart. None of them establishes a result on other repositories, tasks, or harnesses.
- The rank, cost rank, cost shares, frontier steps, and component gaps are aicharts derivations from the snapshots named in each caption. A configuration added, removed, or rescored by Artificial Analysis moves them, and the coding-agent snapshot advances daily.
- Gemini 4 Argon in Cursor, Codex, Claude Code, or any harness other than Antigravity CLI is a configuration this snapshot does not store, so this note says nothing about it. The snapshot also stores no earlier Google model in Antigravity CLI, so there is no same-harness generation step to report.
- Access to the model is limited to the groups Google describes, and Google has published no general release date. A reader outside that group cannot reproduce the row’s cost or score today, and the introductory list price behind the cost column is one Google has already said will rise.
- The two charts use different task sets and different cost definitions. Their scores are not one ranking.
- The prices, the access statements, and the vendor-run benchmark figures belong to Google. aicharts did not run Gemini 4 Argon.
