Skip to notes
AI Charts
Theme
Appearance

Claude Opus 5.5 leads the Intelligence Index at 57.6 for $5.98

Claude Opus 5.5 at max effort scores 57.6 on the Intelligence Index at $5.98 per task, first of 97 configurations. The four highest points on the cost frontier are all its effort levels.

Drafted with AI and reviewed by weekday-monitor AI editorial review.

Five matte ivory spheres of decreasing size rest one per step on a rising charcoal staircase, with a brass line along the top step.
Each effort level of Claude Opus 5.5 is its own row on the Intelligence Index, and the top four steps of the cost frontier belong to the same model. AI Charts editorial illustration · Slopcamera with GPT Image 2

Anthropic released Claude Opus 5.5 on September 22, 2026. In the Intelligence Index snapshot retrieved Sep 23, 2026, 1:38 PM UTC, Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) scores 57.6 at $5.98 per task, first of the 97 comparable configurations, meaning the rows with a measured cost per task, and on the chart’s cost frontier. The four lower effort levels of the same model score 56.0, 53.6, 51.2, and 42.3. Walking the cost frontier down from the top, the first four points are all Claude Opus 5.5 rows. The coding-agent chart stores one configuration running Claude Opus 5.5, scored on three coding benchmarks inside a harness rather than on the Intelligence Index.

One configuration, 10 evaluations

The Artificial Analysis Intelligence Index runs a model through its API under one harness that is the same for every model, across 10 evaluations weighted 30% agents, 20% coding, 20% scientific reasoning, and 30% general capability, at version 4.3.2. The evaluations are AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, and AA-LCR v1.1. Cost per task is the weighted average API bill for one task across those evaluations, split into input-side cost (non-cached input, cache reads, and cache writes) and output-side cost (reasoning and answer tokens).

A row on the chart is one configuration: the model at one effort level with the settings Artificial Analysis names in the row. The Claude Opus 5.5 rows share the phrases Adaptive Reasoning and Default Fallback and differ in effort level. The snapshot records the effort level and does not define the other two phrases, so this note does not interpret them.

First of 97 comparable configurations

The 57.6 places Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) first of the 97 comparable configurations in the snapshot retrieved Sep 23, 2026, 1:38 PM UTC. No configuration scores higher. The row is on the chart’s cost frontier: no comparable configuration scores at least as high at the same or lower cost per task.

No other configuration scores within one index point of it. The nearest score below is Claude Opus 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback) at 56.0, 1.6 points lower for 0.6x the cost per task.

The table lists the five highest-scoring configurations after it, with each cost as a multiple of its $5.98. Two of the five are Claude Opus 5.5 at a lower effort level. The highest-scoring configuration from another model is Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) at 53.4, 4.3 points below at 1.3x the cost per task.

The five highest-scoring comparable configurations after Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) in the snapshot retrieved Sep 23, 2026, 1:38 PM UTC
ConfigurationIntelligence IndexCost per taskPoints below Claude Opus 5.5 (max)Multiple of its cost
Claude Opus 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback)56.0$3.461.6 points0.6x
Claude Opus 5.5 (Adaptive Reasoning, High Effort, Default Fallback)53.6$1.824.0 points0.3x
Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)53.4$7.634.3 points1.3x
Claude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback)53.2$5.984.4 points1.0x
GPT-6 Astra (max)52.7$3.264.9 points0.5x

Walking the cost frontier down from its highest-scoring point, the first four points are Claude Opus 5.5 rows: max, xhigh, high, and medium. The first frontier point from another model is GPT-6 Sol (max) at 47.5 for $1.06 per task, so every frontier point that costs more than $1.06 per task is a Claude Opus 5.5 row.

Five effort levels of one model

The snapshot stores five comparable Claude Opus 5.5 rows with an effort level, one per level, and no row without one. From low to max, the score moves from 42.3 to 57.6 index points and the cost per task from $0.551 to $5.98. The table lists the levels cheapest first and states what each step buys over the level below it.

Comparable Claude Opus 5.5 effort levels in the snapshot retrieved Sep 23, 2026, 1:38 PM UTC, cheapest first
Effort levelIntelligence IndexCost per taskOutput tokens per taskReasoning share of output tokensPoints over the level belowCost multiple of the level below
low42.3$0.55110,15133%--
medium51.2$1.3425,74545%+8.92.4x
high53.6$1.8235,58451%+2.31.4x
xhigh56.0$3.4665,66761%+2.41.9x
max57.6$5.98119,16670%+1.61.7x

The largest step is from low to medium: +8.9 points at 2.4x the cost per task. The last step, from xhigh to max, adds 1.6 points at 1.7x the cost per task. At max, the model writes 119,166 output tokens per task and 70% of them are reasoning tokens; the share runs from 33% to 70% across the five levels.

Anthropic’s announcement names medium as the default effort level. In the snapshot the medium row scores 51.2 for $1.34 per task, 6.4 points below max at 0.2x its cost. The snapshot dates the max row September 22, 2026 and the low, medium, high, and xhigh rows September 17, 2026. The snapshot stores no Opus 5.5 row without an effort level, and Anthropic’s announcement states that the model is no longer offered with thinking switched off.

Where the $5.98 goes

Of the $5.98 per task at max, $3.60 is input-side cost and $2.38 is output. Largest component first, the five parts are cache reads at $2.42, reasoning tokens at $1.68, cache writes at $1.07, answer tokens at $0.705, and non-cached input at $0.103. Cache reads are 41% of the total. Input-side cost is 60% of the total at max and stays between 60% and 63% across the five levels, while the reasoning share of output tokens moves from 33% to 70%.

Cost per task by component for each comparable Claude Opus 5.5 effort level in the snapshot retrieved Sep 23, 2026, 1:38 PM UTC, cheapest first
Effort levelInput costOf which cache readsOutput costOf which reasoning tokensInput share of total
low$0.348$0.120$0.203$0.06863%
medium$0.821$0.410$0.515$0.23461%
high$1.11$0.604$0.712$0.36561%
xhigh$2.15$1.37$1.31$0.80662%
max$3.60$2.42$2.38$1.6860%

Artificial Analysis’s model page, captured September 25, 2026 UTC, lists the prices behind these figures as $4.00 per million input tokens and $20.00 per million output tokens with a 95% cache discount, based on Anthropic’s API, and rounds the row to an index score of 58 at $5.98 per task. It records 260M output tokens to run the whole index, which it calls “very verbose in comparison to the median of 88M,” lists the model as proprietary with a 1M token context window, released September 22, 2026, and sums the row up as “amongst the leading models in intelligence, but somewhat expensive when comparing to other models of similar price.”

Anthropic’s prices and claims

Anthropic’s announcement of September 22, 2026 opens with the claim that “It performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5.” It prices the model at $4 per million input tokens and $20 per million output tokens, against $5 and $25 for Opus 5, and says that “Cache reads (which make up the majority of agentic and coding work costs) are $0.20 per million tokens, 60% less than Opus 5.” In the snapshot’s max row, cache reads are $2.42 of the $5.98 per task. Developers reach the model on the Claude Platform as claude-opus-5-5.

The announcement reports Anthropic’s own runs: Terminal-Bench 4.0 at 66.4% at xhigh effort, plus FrontierCode v1.1 (Main), CursorBench 4.0, GDPval-AA v2.1, AutomationBench, Humanity’s Last Exam, Terminal-Bench-Science 0.1, OSWorld 2.0, and Chartography. None of those figures appears on an AI Charts chart. Anthropic states that “Unless otherwise noted, all Claude Opus 5.5 results use adaptive thinking at max effort,” and adds that “at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences.” Terminal-Bench 4.0 and GDPval-AA v2.1 are also two of the 10 evaluations inside the Intelligence Index, where Artificial Analysis runs them under its own harness; the snapshot stores the composite score, not the per-evaluation results, so the 66.4% cannot be checked against it here.

Anthropic also writes that “Opus 5.5 is the first Opus model to launch with a similar class of safeguards to Fable 5.1 on cybersecurity, biology, and distillation, all of which fall back to another model transparently.” The snapshot names each Opus 5.5 row with the phrase Default Fallback and records nothing else about that setting, and the model page does not define it, so whether the two describe the same behavior is not something either source states.

Claude Code · Opus 5 is a different model on a different chart

The coding-agent chart is a daily snapshot of the public Artificial Analysis coding-agents comparison. A row on it is a model running inside a named agent harness at one effort setting, and its AA Index averages three coding benchmarks: DeepSWE v1.1, Terminal-Bench 4, and SWE-Atlas-QnA. Its cost is the mean API bill for one task in that harness, including every tool call and repeated context the harness sends.

The coding-agent snapshot retrieved Sep 25, 2026, 2:39 PM UTC stores one configuration that runs Claude Opus 5.5: Claude Code · Opus 5.5 (max) at 66.0 on AA Index for $13.04 per task. It also stores the previous generation, Claude Code · Opus 5 (max), at 59.7 for $10.79 per task.

The Intelligence Index row and the coding-agent row that readers most often set side by side, each scored on its own task set with its own cost definition
ChartConfigurationModel generationScoreCost per taskSnapshot retrieved
Intelligence IndexClaude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback)Claude Opus 5.557.6$5.98Sep 23, 2026, 1:38 PM UTC
Coding agents (AA Index)Claude Code · Opus 5 (max)Claude Opus 559.7$10.79Sep 25, 2026, 2:39 PM UTC

The Intelligence Index snapshot retrieved Sep 23, 2026, 1:38 PM UTC stores no Claude Opus 5 row, so the step from Opus 5 to Opus 5.5 that Anthropic describes cannot be measured on the Index from this snapshot.

Limits

  • The scores, costs, and token counts above are Artificial Analysis measurements of the Claude Opus 5.5 rows on the retrieval date under Intelligence Index version 4.3.2, and of Claude Code · Opus 5.5 (max) and Claude Code · Opus 5 (max) under DeepSWE v1.1, Terminal-Bench 4, and SWE-Atlas-QnA. They say nothing about other tasks, prompts, or harnesses.
  • The rank, frontier walk, nearest-score table, effort ladder, and cost shares are computed from those snapshots by AI Charts. A new, removed, or rescored configuration moves them, and both snapshots update on a schedule.
  • The coding-agent scores above are the Claude Code · Opus 5.5 (max) row in the snapshot. They say nothing about Claude Opus 5.5 inside a harness the snapshot does not store.
  • The Adaptive Reasoning and Default Fallback settings in the row names are recorded by Artificial Analysis and not defined in the snapshot; the per-evaluation scores behind the composite are not stored either.
  • The prices, the 40% cost claim against Opus 5, and the benchmark table are Anthropic’s. AI Charts did not run Claude Opus 5.5.

Sources

  1. LLM LeaderboardArtificial Analysis, 2026. The public models leaderboard is the source of the AI Charts Intelligence Index snapshot. Scores, per-task costs, and output tokens are Artificial Analysis measurements under Intelligence Index v4.3.2.
  2. Claude Opus 5.5 (max with fallback) - Intelligence, Performance & Price AnalysisArtificial Analysis, 2026. Cited, from the page captured September 25, 2026 UTC, for the 58 index score, the proprietary label, the September 22, 2026 release date, the $4.00 and $20.00 per million token prices with a 95% cache discount based on Anthropic’s API, the $5.98 cost per index task, the 260M output tokens across the index, the 1M token context window, and its summary and verbosity sentences.
  3. Introducing Claude Opus 5.5Anthropic, 2026. Cited for the September 22, 2026 release, the $4 and $20 per million token prices against $5 and $25 for Opus 5, the $0.20 cache-read price, the 40% cost claim against Opus 5, medium as the default effort level, the max-effort setting behind Anthropic’s benchmark table, the statement that thinking can no longer be switched off, the safeguard fallback sentence, and the vendor-run benchmark table that this site does not chart.
  4. Coding AgentsArtificial Analysis, 2026. The public coding-agents comparison is the source of the AI Charts coding-agent snapshot. Model names, agent harnesses, settings, AA Index scores, and mean API costs are Artificial Analysis measurements.

Figures come from the cited primary sources and the AI Charts datasets. AI Charts did not rerun the reported benchmarks. Reported results apply to the named source, workload, configuration, and observation date. They do not establish performance on every task or product.