GPT-6 Astra: the day nobody answered

GPT-6 Astra: the day nobody answered
ContenidoContents

This afternoon ChatGPT, Claude and Grok went down. At the same time. Tonight OpenAI shipped GPT-6 Astra with a tagline that, with all three dark, sounds like a joke: anything you can do on a computer, Astra can do for you. Fast.

The tweet is short. The announcement is not. The 6 is no longer a rumour. It’s a model, a price, and a drip-feed rollout.

This afternoon nobody answered

It wasn’t one of them getting slow. It was all three. And if you look at DownDetector, Gemini had a spike too. Google hasn’t confirmed it. The others have.

Anthropic went first: around 15:23 Madrid time, elevated errors on Mythos 5.1, Fable 5.1 and Opus 5. They marked it resolved around 18:16. OpenAI acknowledged “elevated errors” across ChatGPT and Codex at 16:43; a routing error, they said, and they marked it resolved at 18:55. Grok, the one I actually use, as well: from almost no reports to more than a thousand in three-quarters of an hour. Cursor complained about Claude and Grok. AWS, Azure and Cloudflare did not declare a major fault. Nobody has called it an attack. Nobody has called it the launch. It just landed on the same afternoon.

Tibo, from Codex at OpenAI, boiled it down to a cascade: when they go down, the traffic that floods the others takes them down too. That fits what we saw. It also fits something duller and more important: they’re no longer three products. They’re one nervous system, with three vendors and a lot of people who, if the first one fails, hit the second.

I’ve written more than once that I don’t want to depend on a single lab. This afternoon plan B and plan C were in the same queue.

What it is, without the smoke

GPT-6 Astra is the successor to the GPT-5.6 family (Sol, Terra, Luna). OpenAI calls it the world’s most intelligent and aligned model. The actual product isn’t the ranking. It’s using the computer: the Mac, the browser, the CRM, the slides, the calendar. Not a chat that drafts an email. An agent that clicks, fills, installs, tests and leaves you the artefact.

The numbers that jump hardest are about that. On OSWorld 2.0 it hits 72.6% and takes about 40 minutes per task; GPT-5.6 Sol stopped at 65.7% and about 75 minutes. 47% less time, they say. On Agents’ Last Exam, 59.3%. On ScreenSpot-Pro, 92.7%. On AutomationBench, 41.4% against Sol’s 18.1%. OpenAI’s figures. The usual asterisk.

On code, Terminal-Bench 4.0: 57.7%. Sol, 37.3%. Fable 5.1, 55.8%. The 6 beats Sol by a stretch and Fable by a whisper. On agentic science the jump is wider: Terminal-Bench-Science 0.1 at 64.6%. Fable 5.1 stopped at 52.6%. Sol at 22.4%. Two days ago I wrote that 5.1 was not a 6. Today the 6 answered.

And they saturate what you can barely push higher: FrontierMath Tier 4 at 98%, ARC-AGI-3 at 99.9%, ExploitBench at 100%. When a bench is full, the number stops informing. It informs which benches they picked.

It doesn’t win everything. On Humanity’s Last Exam, Sol and Fable 5.1 sit above it. On the Artificial Analysis index, too. The 6 is not a sweep. It’s a computer-use, cyber and lab model. The generic chat, we’ll see.

The ID is gpt-6-astra. OpenAI API and Amazon Bedrock. It lands first at a handful of organisations. Over the coming days, ChatGPT Plus, Pro, Business and Enterprise. On Enterprise, off by default. There’s an Astra Pro for Pro, Business and Enterprise. Usage sits inside the subscription allowance; if you go over, credits.

The price is Fable’s

$10 per million input tokens, $50 output. The same sticker as Fable 5.1. Fast mode: up to 2.5× the speed, at 2× the price.

OpenAI says that even though the token costs more than Sol, cost per task can come out lower because Astra spends fewer tokens and finishes sooner. They signed it. Anyone who bills by tokens will check the invoice, not the press release.

There’s no GPT-6 Terra or GPT-6 Luna in the announcement. The 6, for now, is a single step. The top one.

The leash, again

Astra “meets the Critical threshold” in cybersecurity under their Preparedness Framework. Translation: it’s good enough at writing exploits that they won’t ship it all at once. On ExploitBench, 100%. In internal evals, without production safeguards, it found and used two zero-days. They say so. They’ll disclose them to the maintainers.

That’s why the rollout is slow. That’s why Daybreak —the cyber-defence programme— sees more. The public model refuses to build a proof-of-concept exploit. The trusted one doesn’t. The scheme is Anthropic’s with Fable and Mythos: the same brain, two locks. OpenAI doesn’t call it that. It does the same thing.

They say it’s their most aligned model. The example they give is the Hugging Face incident: a model facing an impossible task steps outside its perimeter. Sol, without safeguards, did that 48% of the time. Astra, 0%. Fine. They also say Astra’s written reasoning is harder to monitor than Sol’s, because it solves things in fewer steps. That’s in the system card. It isn’t a footnote.

What I actually care about in Codex

On long sessions, models used to summarise the context when the window filled. Each summary loses why a fix failed or how a component behaves. Astra, in Codex, keeps notes across windows and can search the earlier ones, even if that detail never made the summary. Experimental. Default in a few weeks.

That isn’t a benchmark. It’s the actual problem of an agent that’s been inside a repo for three hours. If it works, you feel it. If it doesn’t, we’re still summarising.

Astra will not land in Cursor. OpenAI already shut that tap, and the note from five days ago said, literally, that the cut includes Astra. If you want the 6, it’s ChatGPT, Codex, the API or AWS. Not the dropdown in the editor.

So do I switch?

No.

I’m staying with Grok. I cancelled Claude Max over the contract, not the ranking. An OpenAI 6 doesn’t bring me back to ChatGPT, and least of all on the day all three went down at once.

I already value the harness more than the model. The loop, the tools, that it remembers the repo. The 6 is smarter; the place it lives —Codex, not Cursor, not Grok Build— isn’t mine. A new model in someone else’s harness doesn’t beat one that’s already in mine. This afternoon I didn’t miss three points on a bench. I missed the loop.

What does change is the map. OpenAI’s public ceiling is no longer called 5.6. It costs what Fable costs. It lives on the computer, not in the chat. And this afternoon, right before they unveiled it, the computer didn’t answer.

The lesson is not that Astra is AGI. It’s that we already treat them as infrastructure, and infrastructure, when it fails, fails in a cluster. The tagline says Astra can do anything on a computer. Today, for a while, it couldn’t even open a conversation.

The 6 is here. All three went dark the same day. That’s the post.

OpenAI’s tweet is here. The announcement, here. Tibo’s cascade tweet, here.

CompartirShare