Claude Sonnet 5.5 is here: near Opus, at Sol's price

Claude Sonnet 5.5 is here: near Opus, at Sol's price
ContenidoContents

Six days ago Anthropic shipped Claude Opus 5.5: Fable 5.1-level on most work, at an Opus price. That same night OpenAI released GPT-6 Sol at $2 per million input tokens and $10 output. Today, 28 September, the second model in the 5.5 family arrives: Claude Sonnet 5.5.

The token price does not move. It stays at $2 and $10, the same as Sonnet 5 and the same as Sol. Anthropic says it generates output more than 30% faster than Sonnet 5 and that, in their own testing, a task costs up to 30% less because it spends fewer tokens. Artificial Analysis, the same day and at max effort, measures a different curve on their index: $7.60 per task, about 50% more than Sonnet 5, with about 193,000 output tokens per task. The most they have recorded.

This is not the verdict of an afternoon inside Claude Code. It’s the spec sheet, read next to someone who didn’t sign it.

The numbers, plainly

  • ID: claude-sonnet-5-5. No date in the identifier. Available today on the Claude API, AWS, Google Cloud, and Azure. Zero data retention, as with Sonnet 5 and Opus 5.5.
  • Context: 1 million tokens. Output: up to 128,000. Knowledge cutoff: June 2026. Same as Opus 5.5.
  • Effort: low, medium, high, xhigh, and max. The API default is high. In the apps and in Claude Code, medium. The levels are recalibrated against Sonnet 5: the same word is not the same amount of thinking.
  • 5-minute cache: reads $0.20, writes $2.50. The 1-hour write is $4. The shortest prompt that can be cached drops from 1,024 tokens to 512.
  • Haiku 5.5 follows in the coming weeks. If inference has to stay in the US, input and output are charged at 1.1×.

In June I wrote that Sonnet 5 would rise to $3 and $15 on 1 September. On 10 August Anthropic made $2 and $10 the permanent rate. The platform pricing page still prints the September step. claude.com/pricing does not: Sonnet 5 and Sonnet 5.5 are both $2 and $10.

Per million tokens. Cache writes are the 5-minute tier. Where I don’t have the figure, the cell stays empty.

Model Input Cache read Cache write Output
Grok 4.7 $2 $0.50 $6
Sonnet 5.5 $2 $0.20 $2.50 $10
GPT-6 Sol $2 $0.20 $2.50 $10
Opus 5.5 $4 $0.20 $5 $20
GPT-6 Astra $10 $1 $50

Sonnet’s input already costs what Grok 4.7’s does. Output does not: $10 against $6. Cache reads run the other way: $0.20 against $0.50, the same price as Opus 5.5 and Sol. On batch, input and output drop by half.

In August I cancelled Max over the watermark, not over the model. This sheet does not put me back on the subscription. It changes what someone still on the API pays. I still write code with Grok: the 4.7 sheet has not moved.

The benchmark, effort first

Anthropic publishes a summary table and does not name the effort of every cell. Where they do, I write it down. Opus 5.5’s 66.4% on Terminal-Bench 4.0 is its score at xhigh. On FrontierCode, Sonnet 5.5 scores 52.1% at xhigh and 46.2% at max. At the highest effort it more often runs the code-review skill, splits the work across subagents, and in two cases Cognition looked at, that ended in a timeout or in edits outside the task’s scope. FrontierCode penalizes both, so the score drops.

Sonnet 5’s 10.3% on Terminal-Bench 4.0 is not the 80.4% I quoted in June. That was Terminal-Bench 2.1. This is a different test.

Benchmark Sonnet 5.5 Sonnet 5 Opus 5.5 GPT-6 Sol
Terminal-Bench 4.0 70.6% 10.3% 66.4% (xhigh) —
FrontierCode 1.1 52.1% (xhigh) / 46.2% (max) 42.4% 54.4% 49.3%
CursorBench 4.0 55.5% 34.1% 57.8% —
GDPval-AA v2.1 (Elo) 1844 1449 1846 1487
AA-Briefcase v1.1 (Elo) 1811 1359 1822 1483
Humanity’s Last Exam, with tools 64.5% 54.9% 67.7% —
OSWorld 2.1, partial 80.1% 57.0% 81.8% —
Chartography, no tools 61.6% 15.6% 64.4% 53.6%

Artificial Analysis ran GDPval-AA and AA-Briefcase on a pre-release deployment. Anthropic later found a structured-output bug that could degrade the reply. They expect the effect, if any, to be small and to leave the score short of the real one. The bug is fixed. On Chartography, OpenAI has just fixed an image-understanding bug in GPT-6 Sol. Anthropic says their internal tests do not move that cell, and that Artificial Analysis does not expect a large change on Briefcase or GDPval. Opus’s 81.8% on OSWorld matches the number from 22 September. That row was called OSWorld 2.0.

On CursorBench, Sonnet 5.5 lands 2.3 points from Opus 5.5. On GDPval-AA, 2 Elo points: 1844 against 1846, and about 400 above Sonnet 5.

What Anthropic wants you to look at is the low effort, not the peak. On Terminal-Bench 4.0, Sonnet 5.5 at medium — the default in the apps — beats Sonnet 5’s best score for less than a tenth of the cost per task. On CursorBench, low already beats that best score, also for less than a tenth. On AA-Briefcase, medium beats it for about a ninth. On FrontierCode, high — the platform default — scores 10 points above Sonnet 5 at the same effort, for about a fifteenth of the cost, and matches Sol’s best score for about a fifth.

At max, on several of those benches, it sits next to Opus 5.5. Anthropic says so and qualifies it in the same paragraph: in their own use, and in that of outside testers, Opus 5.5 remains clearly better at open-ended work that needs sustained judgment. The ceiling of the house is still Fable, at $10 and $50.

Anthropic says this is the first Sonnet to beat Pokémon Red working only from screenshots. That is a launch sentence. I keep it because it is concrete, and because it does not replace the table.

What someone who didn’t sign the post measured

Artificial Analysis places Sonnet 5.5 (max) second on their intelligence index: 56, two points behind Opus 5.5 (max, 58) and 18 ahead of Sonnet 5. On Terminal-Bench 4.0, in their harness, it scores 64%. Sonnet 5 (max) lands 50 points lower. Opus 5.5 and Astra (xhigh) sit around 60%. That is not the 70.6% in Anthropic’s table. A different harness. They stay apart.

To get there it spends about 193,000 output tokens per index task. 60% more than Opus 5.5 (max) or Sonnet 5 (max), and about seven times Astra (max). The bill for that task: $7.60, about 50% above Sonnet 5. At Sol’s price, Artificial Analysis leaves it off the cost frontier of their index. At high effort it is most competitive: a little behind Sol, at practically the same cost per task. At the lower efforts, some Astra or Sol settings deliver equivalent performance for less money.

On factual knowledge, AA-Omniscience, it is right 54% of the time against Opus’s 66%, and it hallucinates less: 47% against 59%. On Humanity’s Last Exam and SciCode it sits about 6 points below Opus. That gap is not the 64.5% with-tools cell in the table above.

Anthropic’s fallback fired on about 0.1% of index tasks, almost all of them on Terminal-Bench 4.0, and always down to Sonnet 5. These runs are on the pre-release deployment, the one with the structured-output bug. Artificial Analysis says the effect should be small, or should leave the score short, and that they will run them again.

What the people who tried it say

Quotes from the launch page. Not a public benchmark.

Balyasny, on 2,441 private finance tasks: ahead of Sonnet 5, at about 121,000 tokens per answer against 497,000. Box: 2.4× faster and 12% fewer tokens overall, and it checks the source document again. Base44, on 118 real app builds: level with Opus 5, at 3.6 iterations on average against 7.7, and the lowest rate of failed tool calls among the models they compared. Unity counts a task when the change works at runtime, not when the model says it is done. Most of Sonnet 5.5’s work passed that check, and it completed 90% of their multi-step benchmark inside the editor.

In an internal test, Anthropic gave it a public company’s earnings materials, the analyst-call transcripts, and a template, and asked for 10 slides. Two experts judged the first draft ready to send. Their test, two judges.

If you code against the API: what breaks

Swapping the ID is not enough. The migration guide is long. This is what returns a 400, or a 200 that goes quiet:

  • thinking: disabled is gone. Send between_tools. It is accepted only at low, medium, and high. At xhigh or max it returns 400. With between_tools, effort cannot change mid-conversation. If server-side fallback drops you to Sonnet 5, that request runs there with thinking off.
  • Forced tool use is gone. tool_choice of any or of a named tool returns 400, including on the token-counting endpoint. auto and none remain. For JSON that matches the schema, use auto and mark the tool strict. At most 20 strict tools. On Bedrock, structured outputs are not available for this model: send auto without strict, and validate yourself.
  • Thinking is tied to the model and the account. Sonnet 5.5 reads blocks from Sonnet 5, Opus 4.8, Haiku 4.5, and earlier models. It does not read blocks from Opus 5, Opus 5.5, Fable, or Mythos. The API drops them, returns 200, and does not bill them. On accounts created on or after 31 August 2026, 00:00 UTC, the block is signed over the conversation so far: edit the history and replay it, and you get a 400. The conversation has to grow only at the end. And the block works only in the account that produced it, or in an account linked to it. Switch accounts mid-session in Claude Code and it drops. They explain it under preserved thinking.
  • The old computer-use tool does not work on the Claude API or on Google Cloud. Move to computer_toolset_20260801. Bedrock still takes computer_20251124. computer_20250124 is rejected everywhere.
  • The advisor tool accepts fewer models. Opus 4.8, Opus 4.7, Opus 4.6, Sonnet 5, and Sonnet 4.6 as advisors return 400. The advice comes back encrypted: the response text does not include it.
  • Longer text between tool calls comes back inside thinking blocks. At the default display those blocks arrive empty. An interface that showed that text as progress goes quiet. Short remarks stay text.

The API’s default effort is high, and Claude Code’s is medium. Copy a Sonnet 5 configuration and you are not at the same point on the curve. Cost has to be measured again.

The leash, now on Sonnet too

Until today the heavy cybersecurity leash belonged to Fable, to Mythos, and, since the 22nd, to Opus 5.5. Sonnet 5.5 is the first Sonnet to ship with safeguards and a fallback of that kind. Anthropic says its cybersecurity capability is comparable to Opus 5’s.

Finding and fixing bugs in your own code stays on this model. Higher-risk cybersecurity tasks are visibly routed to Sonnet 5. The verification program is widening so people can apply for tiered access on Sonnet 5.5, Opus 5.5, and the Mythos models. Biology safeguards are Sonnet 5’s. Most life-sciences work is unchanged. Some microbiology and virology requests may be flagged in error.

It is also the first Sonnet with classifiers against reasoning extraction. A refusal returns stop_reason: refusal and can name the category: cyber, bio, frontier_llm, reasoning_extraction, or general_harms. Server-side fallback, in beta and on the Claude API only, retries cyber and frontier_llm on Sonnet 5. It does not retry the other three.

On the automated audit, about 1,850 scenarios, it improves on or matches Sonnet 5 on most measures of alignment and honesty. On the newer containment evaluations it comes close to Opus 5.5, the best model they tested, in how rarely it tries to leave the sandbox, and it is the least likely of their models to probe the limits of its containers. Across the full audit, Opus 5.5 is still slightly ahead. They say they found no evidence that it pursues goals that conflict with the user’s. And they say, as in the recent assessment, that no suite of tests catches every failure.

What it means

For scoped work — a bug, a document, slides from a template — Anthropic wants you to leave this Sonnet on and keep Opus 5.5 for the open-ended jobs. The price invites that: half the input and output of Opus, the same cache read, and the 5-minute cache write also at half, $2.50 against $5.

The line “up to 30% cheaper” fits on their page. It does not fit on the Artificial Analysis index. Both readings can be true at once. One is measured at the effort they want you to live on: medium in the app and in Claude Code, high on the API, compared with Sonnet 5. The other is measured at max, where this model writes more than any they have put on that scale. If the agent lives at max, the bill can go up. If it lives where Claude Code leaves it by default, Anthropic’s bet runs the other way, and no outside measurement of that cell has been published yet.


Sources: Anthropic — Introducing Claude Sonnet 5.5, model page, migration guide, claude.com/pricing, Artificial Analysis, announcement.

CompartirShare