Claude Sonnet 5.5: Cheap Below Max, Not at the Top

Sonnet 5.5 keeps the $2 and $10 rate and is the cheap option below max effort. At max it can cost more than Opus 5.5 for a similar or lower score.

Claude Sonnet 5.5: Cheap Below Max, Not at the Top

Claude Sonnet 5.5: Second on the Index, First on Tokens

Claude Sonnet 5.5 came out on September 28, 2026, at the same list price as Sonnet 5 and as GPT-6 Sol: $2 per million input tokens and $10 per million output. Anthropic calls it a faster complement to Opus 5.5 for well-scoped work. Artificial Analysis, running it independently at max effort with the default fallback on, puts it second on the Intelligence Index: 56, two points behind Opus 5.5 at 58, and 18 points ahead of Sonnet 5 at 38.

That second place is the expensive end of the model. The same run used about 193,000 output tokens per index task, the highest Artificial Analysis has measured, about 60% more than Opus 5.5 at max and about seven times GPT-6 Astra at max. The task costs $7.60. Opus 5.5 at max costs $5.98. Same family, half the uncached rate card, a higher bill at the setting that produces the headline score.

Intelligence Index bars and the cost-per-task scatter for the Sonnet 5.5 release
Artificial Analysis. Sonnet 5.5 at max is second on the index and sits to the right of the cost frontier, beside Opus 5.5.

The setting is the product

The API id is claude-sonnet-5-5. Context is 1 million tokens. Claude Code and the Claude apps default to medium. The Claude Platform defaults to high. Neither default is the score of 56.

Artificial Analysis ran all five efforts. These are their index scores and the cost of one Intelligence Index task.

Effort Intelligence Index Cost per task Tokens in the full index run
Low 36 $0.41 23 million
Medium 41 $0.59 29 million
High 47 $1.08 50 million
Xhigh 52 $2.74 100 million
Max 56 $7.60 410 million

From high to max, the index moves from 47 to 56 and the task price moves from $1.08 to $7.60. Four of those nine points arrive at xhigh, for $2.74. The last four points, from 52 to 56, are the jump to $7.60 and to 193,000 output tokens a task. A cache read stays $0.20, the same as Opus 5.5, so a run that is mostly cache hits does not get half of Opus either.

Artificial Analysis says the high setting is the competitive one: narrowly behind GPT-6 Sol on the index, at effectively the same cost per task. Sol at max scores 48 and costs about $1.06 on their earlier run. Sonnet at high scores 47 and costs $1.08. Lower efforts fall behind configurations of Sol or Astra that match the score for less money. Max falls off the frontier entirely, to the right of Opus.

Anthropic's own claim is up to 30% less cost per task than Sonnet 5, from fewer tokens, and output more than 30% faster. That is their workloads. On this index, max effort costs about 50% more per task than Sonnet 5, because it talks more. Both statements can be true. They describe different efforts.

Output tokens per index task, and index versus tokens
Artificial Analysis. The tall bar is Sonnet 5.5 at max, about 193,000 output tokens per task. Low, medium, and high sit far to the left.

Where second place is real, and where it is not

On Artificial Analysis's Terminal-Bench 4.0, Sonnet 5.5 at max scores 63.6%. Opus 5.5 is 59.6%. GPT-6 Astra is 59.1%. GPT-6 Sol is 43.9%. Sonnet 5 at max is 14.1%. Their article rounds the leaders to 64% and 60%. This is a lead, on their harness, with fallback left on. They saw the model fall back to Sonnet 5 on about 0.1% of index tasks, mostly inside this bench.

Terminal-Bench-Science is the check. Sonnet 5.5 scores 53.3%. Astra leads at 63.3%, Opus 5.5 is at 59.0%, and Sol is at 30.0%. A terminal lead on professional tasks does not carry over to the science workflows.

Anthropic's launch table is a third number, and it is max effort on their setup: 70.6% on Terminal-Bench 4.0, priced on their cost chart at $12.54 a task. Opus's best on that chart is 66.4% at xhigh, for $7.35. At medium, the Claude Code default, Sonnet is 28.8% at $0.83 and Opus is 57.6% at $2.94. Quote 63.6% and 70.6% as two harnesses.

Terminal-Bench 4.0 and Terminal-Bench-Science
Artificial Analysis. Sonnet 5.5 leads Terminal-Bench 4.0 at 63.6% and sits third on Terminal-Bench-Science at 53.3%.

Knowledge work is the tie. AA-Briefcase is 1811 Elo against Opus at 1822. GDPval-AA is 1844 against 1846. AutomationBench-AA is 71% against 70%. GPT-6 Sol, at the same $2 and $10, is 1487 on GDPval and 1483 on Briefcase. The sticker matches. The deliverable does not.

The gap is factual knowledge. On AA-Omniscience, Sonnet's accuracy is 54% against 66% for Opus, and its index is 32 against 46. Its hallucination rate is lower, 47% against 59%. It is wrong less often among the answers it gives, and it knows less. SciCode is 61% against 67% for Opus. Humanity's Last Exam, on this index, is 55% against 61%. Anthropic's with-tools Humanity's Last Exam is a different setup, 64.5% against 67.7%. GDP.pdf is a tie at 26%.

These Artificial Analysis numbers come from a pre-release deployment. Anthropic found a structured-output bug on that deployment, since fixed, and expects the scores to understate the public model by a little. Artificial Analysis plans to re-run the affected evals.

Component scores across the Intelligence Index
Artificial Analysis. Knowledge-work bars sit with Opus at the top. The separation is AA-Omniscience, further down the figure.

Anthropic's coding curve says the same thing in dollars

Anthropic's accuracy-versus-cost charts are interactive, so the points below are their labels, not a second copy of the Artificial Analysis figures.

On FrontierCode, high effort, the Platform default, Sonnet scores 49.4% at $0.42. That matches GPT-6 Sol's best, 49.3% at $2.07. Xhigh is Sonnet's best on this bench, 52.1% at $1.59. Max drops to 46.2% and costs $20.78. Opus at medium is already 54.6% at $0.80. The max drop is explained in their footnote: FrontierCode penalizes edits outside the task, and at max the model more often split a code review across many subagents. In two cases Cognition looked at, that ran out of time or edited past the task.

On CursorBench, Sonnet's best is 55.5% at $9.67. Opus at high is already 56.0% at $3.97.

What you have to change in the agent

Sonnet 5.5 is the first Sonnet with cyber safeguards of the kind used on Anthropic's most capable models. Higher-risk cyber tasks fall back to Sonnet 5. Biology safeguards stay at Sonnet 5's set. A refusal is HTTP 200 with stop_reason: "refusal".

Thinking cannot be sent as disabled. The lowest setting is between_tools, and only at high effort or below. Forced tool choice returns an error. Thinking blocks from this model stay inside this model: Opus 5, Opus 5.5, Fable, and Mythos do not read them, and Sonnet 5.5 does not read theirs. On API accounts created on or after August 31, 2026, editing the history before one of its thinking blocks returns a 400. On the Claude API and Google Cloud, the older computer-use tool computer_20251124 is rejected. The advisor tool rejects Opus 4.8, Opus 4.7, and Sonnet 5 as advisors.

In practice

- Treat the score of 56 as max effort. It costs $7.60 an index task and about 193,000 output tokens. Opus 5.5 at max scores 58 and costs $5.98. Buy max when the last four index points are worth that jump, from xhigh at $2.74.

- For the default agent, use high. The Platform already does. The index is 47 at $1.08, next to GPT-6 Sol at max, 48 at about $1.06. Medium, the Claude Code and apps default, is 41 at $0.59, and Artificial Analysis puts Sol and Astra ahead of it on the cost frontier.

- Move well-scoped knowledge work here when the comparison is Opus's uncached rate. Briefcase is 1811 against 1822 and GDPval is 1844 against 1846. Keep open-ended factual work on Opus. Omniscience accuracy is 54% against 66%.

- Keep science-terminal and FrontierCode-at-max off this model. Terminal-Bench-Science is 53.3%, behind Astra at 63.3% and Opus at 59.0%. FrontierCode at max is 46.2% at $20.78. Xhigh is 52.1% at $1.59.

- Quote Terminal-Bench as two results. Artificial Analysis has 63.6%, ahead of Opus at 59.6%. Anthropic's launch chart has 70.6% at max for $12.54, against Opus at 66.4% for $7.35.

- A run that is mostly cache hits costs the same $0.20 read as Opus. The half-price is the uncached input and output.

- Before pointing a Sonnet 5 agent at this id, switch thinking to between_tools, drop forced tool choice, and drop Opus 4.8, Opus 4.7, and Sonnet 5 as advisors. A cyber refusal can fall back to Sonnet 5. A biology refusal stays refused.

Sources: Artificial Analysis on Sonnet 5.5, the effort pages, the comparison with Opus 5.5, and Anthropic's launch. Index, tokens, and the four charts are Artificial Analysis at max with fallback, unless the effort table says otherwise. The FrontierCode and CursorBench prices are Anthropic's charts.

Keep reading

Explore more product news and best practices for teams building with Plataforma Tess prod.

Build with TESS

Turn ideas from this article into working AI workflows.

Create agents, automations, and knowledge-powered workflows in one platform built for teams.