Fable 5.1: smarter, cheaper, more leash
ContenidoContents
Today Anthropic shipped Claude Fable 5.1 and Mythos 5.1. The headline is the usual one: the most capable models for coding and long-running work. The second paragraph, the one you have to scroll for: they are the same model, with two levels of lock. Fable is the public one. Mythos stays behind a trusted-access programme, for cyber defence and the life sciences.
In June I wrote that Fable 5 came with a leash. Three days later they switched it off. On 1 July they switched it back on, with an extra link. Today it’s 5.1’s turn.
Two weeks ago I cancelled Claude Max. Not because of the model. A 5.1 doesn’t bring me back. But the chapter fits: the model gets better. So does the contract.
What it is, without the smoke
Fable 5.1 is the successor to Fable 5. Same Mythos class, same list price: $10 per million input tokens and $50 output. A million tokens of context, 128k of output. The ID is claude-fable-5-1. It’s on Claude.ai, the API, AWS, Google Cloud and Microsoft Foundry. Pro, Max, Team and Enterprise.
Mythos 5.1 is the same brain with more permissive safeguards. For now, a set of US organisations. Everyone else waits on the Cyber Verification Program and the life-sciences one, the latter with the US government in the loop.
Translation: they’ve given us the fast car with the limiter on again. The difference is the limiter is finer now, and the tank costs less to fill.
What actually moves
The numbers that jump hardest aren’t coding. They’re agentic science. On Terminal-Bench-Science 0.1, Fable 5.1 hits 52.6%. Fable 5 stopped at 24.7%. Opus 5 at 29.0%. GPT-5.6 Sol at 22.4%. More than double its predecessor, with the usual asterisk: these are Anthropic’s figures, with production safeguards on.
On Terminal-Bench 4.0 the gap is narrower: 42.0% → 55.8% (Fable) and 60.9% (Mythos). The spread between the two is, they say, the tasks where Fable’s cyber lock steps in. On CursorBench 3.2, 73.4% against Fable 5’s 70.5%.
The announcement brings three lab anecdotes I won’t recite as miracles, and I won’t pretend they aren’t there either: protein binders with a hit rate near 50% (10–15% is typical), a higher-resolution Venus map from Magellan data, and GPU kernels that speed up genomics models by up to 2.5×. They signed it. Some of it was validated outside. We’ll see what holds when someone who didn’t write the press release looks.
On code, Cognition says that on day one they move Opus 5 traffic in Devin to Fable 5.1: same or better result, lower cost per task. CodeRabbit, against the official photo, qualifies it: almost the same bug coverage, fewer nitpicks, and almost 50% slower in review. That fits. 5.1 is not a 6. It’s a Fable that thinks more and talks less.
The real price
The sticker hasn’t moved. What they cut is what a long agent actually pays: cache reads. From $1 per million to $0.25. 75% less.
Anthropic estimates about 25% less on typical workloads and up to 45% on very agentic work. That’s the product. Fable 5 was twice as expensive as Opus and you used it “when it really paid off”. If cache falls to a quarter, the top tier stops being a special-occasion luxury and starts looking like the day-to-day of anyone already living in tool loops.
At low or medium effort, they say, Fable 5.1 matches or beats Fable 5 for less money. The dial isn’t new —Opus 5 already brought it— but here it’s the trick: you don’t need Claude Code’s high for everything. The default is high in Claude Code and medium in Cowork and claude.ai. If you don’t touch it, you pay the expensive mode.
The new leash
It hasn’t come back looser. It’s come back more precise.
On cyber, false positives drop 60%. In part because Fable 5.1 can now find vulnerabilities in source code. Not exploit them. Pentesting, exploit generation and binary analysis still go to Opus. On biology, false positives on benign questions had already dropped 85%; real R&D stays out, unless you’re on Mythos.
The 30-day retention is still the toll. Enterprise Frontier Safeguards promises zero-data-retention-grade privacy with data on your cloud, not Anthropic’s. It rolls out in phases this autumn. Until then, ZDR only for whoever is already eligible. Everyone else: a month of traffic stored “just in case”.
And two new links, the ones that matter to me.
The watermark. Fable 5.1 is among the first models to ship after 2 August, when Article 50 of the AI Act already applies. Text carries a statistical filigree. Files, C2PA. I already explained why that looks like a dumb idea and why I left. It hasn’t gone. It’s been normalised: it now ships with the top model.
The thinking lock. On new accounts, you can’t edit an earlier turn and keep the reasoning. The thinking block is bound to the model that produced it and to the conversation prefix. If you touch the history, it drops or you get a 400. Anthropic sells it as defence against distillation: nobody copying the thinking at industrial scale. In the watermark post I already said that industrial agenda travels dressed up as transparency. Here it doesn’t bother dressing up. It’s an API that makes you treat the conversation as append-only.
Three more breaking changes, for anyone still calling Fable 5: forced tool_choice (any / tool) is now a 400; an earlier model can’t read 5.1’s thinking blocks; and editing the past invalidates them. Migrating is not just changing the ID.
So do I go back to Max?
No.
Claude is still very good at writing code. This 5.1, on paper, codes better and costs less to run. Cognition puts it in production on day one. Jane Street says it stays readable on long tasks. Millennium talks about a one-in-a-million crash nobody had explained in years. I believe that enough not to disagree for sport.
I don’t go back because I didn’t leave over the ranking. I left because the product I buy includes a layer I didn’t ask for and can’t turn off. 5.1 doesn’t remove it. It shows it: watermark by default, thinking bound, a month of retention, fallback to Opus when the classifier gets jumpy, and a more capable twin that only the Washington filter gets to see.
In June I said the most dangerous tool isn’t the one that fails, it’s the one you let become indispensable. In July, that you don’t hold the switch. In August, that the floor of a flat is the latest tweet. Today 5.1 is consistent with all three: smarter, cheaper, more leash.
I’m staying with Grok. We’ll see.