Claude Opus 5.5 is here: Fable, at Opus prices
ContenidoContents
Two months ago I wrote that Claude Opus 5 had become the reasonable default: close to Fable, at the then-current Opus price. Today, 22 September, Anthropic ships Claude Opus 5.5 and turns the same screw again. They say that on most work it performs at the level of Fable 5.1 and that running it costs 40% less than Opus 5.
This is not the verdict of a week inside Moodle. It’s the spec sheet, read slowly.
The numbers, plainly
- Model ID:
claude-opus-5-5. - Price: $4 per million input tokens and $20 output. 20% below Opus 5 ($5/$25) and under half of Fable 5.1 ($10/$50).
- Cache reads: $0.20 per million. Opus 5 charged $0.50. In an agent, which rereads the same context over and over, that discount outweighs the cut on a fresh token.
- Context: 1 million tokens. Output: up to 128,000.
- Knowledge cutoff: June 2026.
- Default effort:
medium. On Opus 5, saying nothing meanthigh.
Fast mode, in preview and only on the Claude API and in Claude Code, promises up to 2.5× the speed. It costs double: $8 and $40. Sonnet 5.5 and Haiku 5.5 follow in the coming weeks. Five-hour limits go up on Pro, Max and Team, and subscribers get a rate-limit reset they can save and spend whenever they like.
The benchmark, asterisk first
Anthropic publishes the table at max effort, except where they say otherwise. On Terminal-Bench 4.0, Opus 5.5 runs at xhigh and GPT-6 Astra at high: each model’s own best, as measured by whoever ran it. Production safeguards were on. When they fired, cybersecurity tasks were finished by Opus 4.8, and biology tasks by Opus 5. That pulls the score down on those benches.
| Benchmark | Opus 5.5 | Fable 5.1 | Opus 5 | GPT-6 Astra | GPT-5.6 Sol |
|---|---|---|---|---|---|
| Terminal-Bench 4.0 | 66.4% | 55.8% | 52.3% | 57.9% | 37.3% |
| FrontierCode v1.1 | 54.4% | 50.3% | 48.0% | 53.3% | 47.5% |
| CursorBench 4.0 | 57.8% | 51.8% | 46.6% | — | 41.7% |
| GDPval-AA v2.1 (Elo) | 1846 | 1735 | 1708 | 1542 | 1588 |
| Terminal-Bench Science 0.1 | 58.7% | 52.6% | 29.0% | 64.6% | 22.4% |
| OSWorld 2.0 | 81.8% | 80.7% | 74.0% | — | — |
What I care about is not the peak. It’s the setting I would actually leave on. On FrontierCode, at the default effort (medium), Anthropic says it scores 54.6% and beats Astra’s best (53.3%) at about a fifth of the cost per task. On CursorBench, also at medium, 52.5%: above Fable 5.1’s max (51.8%) and Opus 5’s max (46.6%). On Terminal-Bench 4.0, medium already beats Opus 5 at max for about a fifth of the cost, and matches Astra at about 40% of Astra’s cost.
On terminal science, Astra is still ahead: 64.6% to 58.7%. Anthropic says so on the same page: at this level, benchmark margins are a poor guide to real differences, and in their own use the gap with Fable 5.1 is narrower than the table suggests.
If you code against the API: four things that break
Changing the model ID is not enough. The what’s-new guide lists four changes that return an error, or that leave the response silent:
- Thinking cannot be turned off.
thinking: disabledand a manual budget (type: enabledwithbudget_tokens) return a 400. Omit the field, or sendadaptive. Effort is the dial. - Forced tool use is gone.
tool_choiceset toanyor to a named tool also returns a 400.autoandnoneremain. For valid JSON, useautowith strict tools, or structured outputs. - Thinking blocks are tied to the model. Opus 5.5 reads blocks from Opus 5 and from earlier Sonnet and Haiku models. It does not read Fable or Mythos blocks. Switch mid-conversation to any model other than Fable 5.1 or Mythos 5.1 on the Claude API, and the previous reasoning is dropped. The request does not fail. It simply continues without it.
- The old computer-use tool,
computer_20251124, is rejected on the Claude API and on Google Cloud. Move to thecomputer_toolset_20260801toolset. On Bedrock the old tool still works.
And one that does not error, which is why it is worse: the notes the model writes between tool calls come back inside thinking blocks. At the default display, an interface that streamed that text as progress goes quiet between tools.
The default effort drops from high to medium. Copy an Opus 5 config without looking and you are not comparing the same point on the curve. At a given effort it also thinks more per turn than Opus 5, most of all at xhigh and max. Measure again. Don’t inherit the number.
The leash, this time on Opus
Until now the heavy leash belonged to Fable and Mythos. I wrote about it when they switched Fable 5 off and when they switched it back on. Opus 5.5 is the first Opus to ship with that class of safeguard on cybersecurity, biology and distillation.
In practice: finding and fixing bugs in your own code still runs on this model. Most cybersecurity tasks are routed to Opus 4.8. Serious biology means applying to the life-sciences verification program. Anthropic says it tries to step outside its boundaries about 85% less often than Opus 5 or Mythos 5.1, and that when it does the attempt is low-severity and self-reported. They also say it often suspects it is being evaluated. That, in their own words, is a hole in the evaluation.
Data retention stays at zero, as on previous Opus models. Not as on Fable.
What it means
For me the reading is July’s, one step down on the bill. Opus 5 was already the model you could leave on. This is the same slot, cheaper per task, faster to generate, and less swollen in prose: the thing people most held against Opus 5. Factory says that at medium it matches Opus 5 at high while using 20–25% fewer output tokens. Box reports a third of the tokens and answers 40% less verbose.
The open question is the leash. If your work is ordinary code, you will not feel it. If your work brushes offensive security or biology, this Opus is no longer July’s Opus: you get sent to the model below, or onto an access list. And if you integrate the API, the detail that hurts more than the price is that the default effort has dropped a notch, and that turning thinking off is no longer an option.
Sources: Anthropic — Introducing Claude Opus 5.5, What’s new in Claude Opus 5.5, Reuters.