Skip to notes
aicharts
Theme
Appearance

Gemini 4 Argon at 63.8: every higher row costs at least 2.2x

Antigravity CLI · Gemini 4 Argon (default) scores 63.8 on AA Index at $5.84 per task, third of 31 configurations. Every row above it costs at least 2.2x as much.

Drafted with AI and reviewed by Claude Fable 5.1.

A matte ivory column stands behind a sheer dark veil near the left end of a long low charcoal shelf; two taller dark pillars stand together at the right end, with one brass line along the shelf’s front edge.
Antigravity CLI · Gemini 4 Argon is third on the coding-agent chart, and every row above it costs at least twice as much per task; Google has limited access to the model to named groups. Made with SlopCamera

In the aicharts coding-agent snapshot retrieved Oct 5, 2026, 2:10 PM UTC, Antigravity CLI · Gemini 4 Argon (default) scores 63.8 on AA Index at a mean API cost of $5.84 per task, third of the 31 configurations that carry an index. Two configurations score higher, and the cheaper of them costs 2.2x as much per task. Google announced Gemini 4 Argon on September 30, 2026 as a model “rolling out to a set of trusted cyber defenders through our Fairwind Program,” not as a general release. The post describes access in terms of those defenders, trusted testers, Google’s own internal teams, and the U.S. government’s voluntary pre-release process, gives no date for wider access, and Artificial Analysis’s model page marks the model “Not publicly available” on its charts. The chart row this note reads is an independent measurement of a model most readers cannot yet buy. The snapshot’s update log records the row on October 5, 2026.

One model inside one harness, scored on three benchmarks

The coding-agent chart is a daily snapshot of the public Artificial Analysis coding-agents comparison. Each row is one configuration: a model, the agent harness that ran it, and an effort setting. Artificial Analysis runs the harness on tasks from three coding benchmarks, DeepSWE v1.1, Terminal-Bench 4, and SWE-Atlas-QnA, and AA Index is the mean of the three scores. The cost is the mean API bill for one task in that harness at list prices, including every tool call and every repeated read of the repository the harness sends.

Gemini 4 Argon’s row is the model inside Antigravity CLI, Google’s coding agent, at the default setting. One task in that configuration used 13.7 million tokens and took 35 minutes of harness time on average. The snapshot stores one configuration running Gemini 4 Argon, so there is no effort ladder to climb and no second harness to compare; the same model at another setting or in another harness would be another row. The other costed Google row on the snapshot is Antigravity SDK · Gemini 3.8 Flash (high) at 41.9 for $2.47: a different model in a different harness, so the snapshot offers no same-harness predecessor for Argon.

Third of 31 configurations

The 63.8 places Antigravity CLI · Gemini 4 Argon (default) third of the 31 configurations that carry an AA Index. Two configurations score higher: Claude Code · Sonnet 5.5 (max) at 68.4 and Claude Code · Opus 5.5 (max) at 66.0. The leader, Claude Code · Sonnet 5.5 (max), is 4.6 points above it at $14.19 per task. Its $5.84 per task is the ninth highest cost of the 31 configurations that carry a cost; six configurations that score lower cost more per task.

The nearest score below it is Codex · GPT-6.1 Sol (xhigh) at 62.9, 0.8 points lower for 18% of the Argon row’s cost. The table lists the four highest-scoring configurations after it, with each cost as a share of its $5.84.

The four highest-scoring configurations after Antigravity CLI · Gemini 4 Argon (default) in the snapshot retrieved Oct 5, 2026, 2:10 PM UTC
ConfigurationSettingAA IndexCost per taskPoints below ArgonShare of Argon’s cost
Codex · GPT-6.1 Solxhigh62.9$1.040.8 points18%
Claude Code · Sonnet 5.5xhigh62.9$3.330.9 points57%
Claude Code · Fable 5.1 (with fallback)max62.2$12.391.5 points212%
Devin Fusion CLI · Claude Fable 5.1 XHigh + SWE-2 Mediumdefault61.7$7.902.1 points135%

Every higher row costs at least twice as much

The chart’s horizontal axis is cost per task, so the question for a row that is not first is what the rows above it charge for their extra points. Every configuration that scores above Argon costs more per task. Claude Code · Sonnet 5.5 (max) charges $14.19 for 4.6 points more, so the $5.84 is 41% of its cost. Claude Code · Opus 5.5 (max) charges $13.04 for 2.2 points more, so the $5.84 is 45% of its cost. The cheapest way to buy a higher score on this snapshot is Claude Code · Opus 5.5 (max) at 2.2x the Argon row’s cost.

Configurations above Antigravity CLI · Gemini 4 Argon (default) in the snapshot retrieved Oct 5, 2026, 2:10 PM UTC, highest AA Index first
ConfigurationSettingAA IndexCost per taskPoints above ArgonArgon’s cost as a share of theirs
Claude Code · Sonnet 5.5max68.4$14.194.6 points41%
Claude Code · Opus 5.5max66.0$13.042.2 points45%

A reader who wants a score above 63.8 on this snapshot pays at least 2.2x the Argon row’s cost per task to get it. That is the row’s position in one sentence: not the highest score, and at most half the cost of any row that scores above it.

The frontier row beneath it gives up less than a point

The chart’s cost frontier contains configurations for which no other row costs no more and scores at least as high, with a strict improvement in either measure. Antigravity CLI · Gemini 4 Argon (default) is on it: nothing on the snapshot scores as high for the same money or less. The question the frontier answers is what a reader gives up by stepping down from this score to a cheaper row that nothing dominates.

The first step down is Codex · GPT-6.1 Sol (xhigh): 0.8 points lower for 18% of the $5.84. That is the step a cost-minded reader will notice: less than a point of index for less than a quarter of the cost. The Argon row’s claim on the frontier is the score itself, not a cost advantage over the row beneath it. The frontier ends at Codex · DeepSeek V4 Flash 0731 (max): 25.0 points lower for 1.5% of the cost.

Cost-frontier configurations below Antigravity CLI · Gemini 4 Argon (default) in the snapshot retrieved Oct 5, 2026, 2:10 PM UTC, highest AA Index first
ConfigurationSettingAA IndexCost per taskPoints below ArgonShare of Argon’s cost
Codex · GPT-6.1 Solxhigh62.9$1.040.8 points18%
Codex · GPT-6.1 Solmedium61.4$0.7052.3 points12%
Codex · GPT-6.1 Sollow57.2$0.4996.5 points8.5%
Codex · GPT-5.6 Lunamax43.2$0.43820.5 points7.5%
Codex · DeepSeek V4 Pro 0813max43.1$0.23820.7 points4.1%
Codex · GPT-6 Lunamax41.1$0.17622.7 points3.0%
Codex · DeepSeek V4 Flash 0731max38.7$0.08525.0 points1.5%

DeepSWE v1.1 carries the composite

AA Index is the mean of three component benchmarks, each ranked here among the configurations that carry it, and the Argon row leads one of them. On DeepSWE v1.1 it scores 78.8, first of 31; on Terminal-Bench 4 it scores 56.1, fifth of 31, tied with Devin Fusion CLI · Claude Fable 5.1 XHigh + SWE-2 Medium (default); and on SWE-Atlas-QnA it scores 56.5, 21st of 31, tied with Claude Code · Sonnet 5.5 (high).

It leads DeepSWE v1.1 by 5.6 points over Muse Code · Muse Spark 1.3 (xhigh) and Codex · GPT-6.1 Sol (xhigh), tied. Claude Code · Sonnet 5.5 (max) scores 10.1 points higher on Terminal-Bench 4 and 10.5 points higher on SWE-Atlas-QnA. The third place is a DeepSWE v1.1 result: the row is fifth of 31 on Terminal-Bench 4 and 21st of 31 on SWE-Atlas-QnA, and the composite rests on the component it leads.

Antigravity CLI · Gemini 4 Argon (default) on each AA Index component in the snapshot retrieved Oct 5, 2026, 2:10 PM UTC, against the best other configuration that carries the component
ComponentArgon scoreRankBest other configurationGap
DeepSWE v1.178.81 of 31Muse Code · Muse Spark 1.3 (xhigh) at 73.2+5.6 points
Terminal-Bench 456.15 of 31Claude Code · Sonnet 5.5 (max) at 66.2−10.1 points
SWE-Atlas-QnA56.521 of 31Claude Code · Sonnet 5.5 (max) at 66.9−10.5 points

Two configurations with a lower AA Index score higher on Terminal-Bench 4: Claude Code · Sonnet 5.5 (xhigh) at 58.1 and Claude Code · Fable 5.1 (with fallback) (max) at 57.6. A reader who cares only about terminal work would order these rows differently from the composite.

The Index row is a different unit

The Artificial Analysis Intelligence Index runs the model through its API under one harness that is the same for every model, across 10 evaluations at version 4.3.2, and its cost per task is the average bill for one of those evaluation tasks. In the snapshot retrieved Oct 6, 2026, 11:15 AM UTC, Gemini 4 Argon (High) scores 52.6 at $1.99 per task, eighth of 103 configurations with a measured cost.

Gemini 4 Argon in the two aicharts snapshots, each on its own task set with its own cost definition
ChartConfigurationScoreRankCost per taskSnapshot retrieved
Coding agents (AA Index)Antigravity CLI · Gemini 4 Argon (default)63.83 of 31$5.84Oct 5, 2026, 2:10 PM UTC
Intelligence IndexGemini 4 Argon (High)52.68 of 103$1.99Oct 6, 2026, 11:15 AM UTC

The 63.8 is a mean of three coding benchmarks run inside Antigravity CLI, and the 52.6 is a weighted average of 10 evaluations run through the API. The $5.84 is the harness’s full bill for one task, 13.7 million tokens per task in this snapshot; the $1.99 is the average bill for one evaluation task under Artificial Analysis’s standardized harness, with 61,558 output tokens per task. The model places third on the coding-agent chart and eighth on the Index. The two ranks come from different task sets, different cohorts, and different cost definitions. Neither rank corrects the other. Artificial Analysis’s model page, captured October 6, 2026 UTC, lists the list prices behind both figures, $2.00 per million input tokens and $10.00 per million output tokens with a 95% cache discount, a rounded index of 53 (#8 of 225 models), and 110M output tokens to run the whole index. The page’s FAQ dates the release to September 30, 2026.

Who can run the model, and at what price

Google’s announcement, titled “Gemini 4 Argon: our next era of frontier intelligence,” describes a model “rolling out to a set of trusted cyber defenders through our Fairwind Program.” It states that “Safely releasing frontier capabilities at this level requires a phased approach” and that Google will keep gathering feedback from early testers “before making Argon available to developers, enterprises, and consumers as soon as possible,” with that wider release “starting with paid API customers and Google AI Ultra subscribers.” No date is given. The same page says that “For trusted defenders and our own internal teams at Google, we’ll be releasing Argon without cyber guardrails.” A reader of the chart should hold both facts at once: the score is an independent measurement, and the configuration it measures is one that Artificial Analysis could reach and most readers cannot.

The announcement sets an introductory price of $2 per million input tokens and $10 per million output tokens, with “cached input tokens priced at 95% off input token price,” and its footnote states that after the introductory period “the price of $4 per 1M input tokens and $20 per 1M output tokens will apply.” The chart’s cost column is computed at list prices, so a list-price change moves the row horizontally while the score stays where it is. The same page raises the output token limit to “an industry-leading 1M tokens, up from the previous 64K tokens.” Artificial Analysis’s model page describes the model as “amongst the leading models in intelligence and reasonably priced when comparing to other models of similar price” with a 1M-token context window. The same page marks the model “Not publicly available” on its charts while its FAQ says the model is “available via API through 1 provider.” The two lines describe one situation: an API exists, Artificial Analysis measured the model through it, and Google’s post says who may use it.

Google’s own figures

Google reports its own DeepSWE v1.1 run: “It sets a new state of the art on DeepSWE v1.1 (77.9%).” Google does not say which harness, task sample, or date produced that figure, so it and the chart’s score are not one measurement taken twice. The chart’s DeepSWE v1.1 score for the Antigravity CLI row is 78.8, measured by Artificial Analysis inside Antigravity CLI on the retrieval date. The announcement also prints vendor-run results on the Vals Index, AutomationBench, CWE-bench v1, and LVBench, none of which appears on the coding-agent chart. AutomationBench-AA is Artificial Analysis’s own variant of one of them and counts toward the Index row above, not toward the coding-agent row. Read Google’s figures as the vendor’s description of its model and read the chart for an independent measurement of one named configuration.

Limits

  • The chart scores, task costs, token counts, and durations are Artificial Analysis measurements of the named configuration on the retrieval date, under DeepSWE v1.1, Terminal-Bench 4, and SWE-Atlas-QnA for the coding-agent chart and Intelligence Index version 4.3.2 for the capability chart. None of them establishes a result on other repositories, tasks, or harnesses.
  • The rank, cost rank, cost shares, frontier steps, and component gaps are aicharts derivations from the snapshots named in each caption. A configuration added, removed, or rescored by Artificial Analysis moves them, and the coding-agent snapshot advances daily.
  • Gemini 4 Argon in Cursor, Codex, Claude Code, or any harness other than Antigravity CLI is a configuration this snapshot does not store, so this note says nothing about it. The snapshot also stores no earlier Google model in Antigravity CLI, so there is no same-harness generation step to report.
  • Access to the model is limited to the groups Google describes, and Google has published no general release date. A reader outside that group cannot reproduce the row’s cost or score today, and the introductory list price behind the cost column is one Google has already said will rise.
  • The two charts use different task sets and different cost definitions. Their scores are not one ranking.
  • The prices, the access statements, and the vendor-run benchmark figures belong to Google. aicharts did not run Gemini 4 Argon.

Sources

  1. Coding AgentsArtificial Analysis, 2026. The public coding-agents comparison is the source of the aicharts coding-agent snapshot. Model names, agent harnesses, settings, AA Index scores, and mean API costs are Artificial Analysis measurements.
  2. Gemini 4 Argon: our next era of frontier intelligenceGoogle, 2026. Cited for the September 30, 2026 announcement, the rollout to trusted cyber defenders through the Fairwind Program, the phased-release and wider-release statements, the introductory $2 and $10 per million token prices with the 95% cached-input discount and the later $4 and $20 prices, the 1M output token limit, the statement that trusted defenders receive the model without cyber guardrails, and the vendor-run DeepSWE v1.1 and other vendor benchmark figures that this site does not chart.
  3. LLM LeaderboardArtificial Analysis, 2026. The public models leaderboard is the source of the aicharts Intelligence Index snapshot. Scores, per-task costs, and output tokens are Artificial Analysis measurements under Intelligence Index v4.3.2.
  4. Gemini 4 Argon (High) Intelligence, Performance & Price AnalysisArtificial Analysis, 2026. Cited, from the page captured October 6, 2026 UTC, for the rounded 53 index score and #8 of 225 rank, the proprietary label, the “Not publicly available” marker, the September 30, 2026 release date, the $2.00 and $10.00 per million token prices with a 95% cache discount, the $1.99 cost per index task, the 110M output tokens across the index, and the 1M token context window.

Figures come from the cited primary sources and the aicharts datasets. aicharts did not rerun the reported benchmarks.