GPT-6 Sol and Luna are here: the 6, at half the price

GPT-6 Sol and Luna are here: the 6, at half the price
ContenidoContents

This morning Anthropic shipped Claude Opus 5.5: it performs like Fable and costs like Opus. Tonight OpenAI adds two GPT-6 models below Astra. Sol is $2 per million input tokens and $10 output. Luna is $0.10 and $0.50. That is half what Sol and Luna cost in the 5.6 generation, at the promotional prices they had. Astra has not moved: it is still $10 and $50.

Three weeks ago GPT-6 was a single model, Astra, at Fable’s price. Today there are three: Astra, Sol and Luna. Terra, the middle model of 5.6, has no version 6.

This is not the verdict of an afternoon inside Codex. It’s the spec sheet, read next to this morning’s.

The numbers, plainly

  • IDs: gpt-6-sol and gpt-6-luna.
  • Context: 1,050,000 tokens. Output: up to 128,000.
  • Knowledge cutoff: Sol, 20 April 2026. Luna, 18 May. Astra stopped at 30 April. The cheap one reaches a month further.
  • Effort: none, low, medium (the default), high, xhigh and max. The charts in the announcement are not at medium.
  • The 50% is measured against 5.6’s promotional price. On the price list, the 5.6 Sol promo is still listed at least through 21 November 2026.

Per million tokens, with the prompt under 272,000 input tokens:

Model Input Cached Output
GPT-6 Luna $0.10 $0.01 $0.50
Grok 4.7 $2 $0.50 $6
GPT-6 Sol $2 $0.20 $10
Opus 5.5 $4 $0.20 $20
GPT-5.6 Sol, promo $4 $0.40 $20
GPT-6 Astra $10 $1 $50

Luna’s output does not fall by half: from $1.20 to $0.50. The input does, from $0.20 to $0.10. The headline rounds. The table doesn’t.

If the prompt crosses 272,000 tokens, the whole request goes to double the input and cache rates, and 1.5× the output. The same kind of cliff I wrote about yesterday for Grok: there it starts at 200,000, and the output doubles too. Here the cut sits further out and the output rises less. It is still a cut on the entire request, not on the slice that crossed the line.

A cache write costs 1.25× the input. Fast mode is double: Sol in fast mode, on short context, is $4 and $20, this morning’s Opus 5.5 sticker. Batch and Flex are half of standard. EU data residency is available only on standard processing.

Sol’s input already costs what Grok 4.7 charges. The output doesn’t: $10 against $6. The cache runs the other way: $0.20 against $0.50, the same cache-read price as Opus 5.5, with a fresh token at half.

What OpenAI says, with the dial showing

They land today in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise and Edu. Not in regular Chat yet. Free and Go users can try Luna in the desktop app. On the API, now. The ChatGPT rollout is gradual, through the day. Tibo, from Codex, also loads a banked usage reset into Plus, Pro and Business accounts. That is an extra for today. What stretches the subscription is that the token counts for half.

The benches are theirs, and they mix effort levels. Competitor scores come from public reports. What they measure may not match production ChatGPT.

On AutomationBench 1.0.6 — business workflows, 47 tools — Sol at xhigh scores 33.2% at $0.27 per task. Astra at low, 30.3%, at 3.9× that cost. Opus 5 at max, 26.9%, at 11.1×. Fable 5.1 with an Opus 5 fallback, at max, 31.4%, at more than 8.9×: the fallback fired on about 40% of tasks, and that money is not in the figure. Sol at its high effort beats the cheap Astra. That says Astra at low loses to Sol at xhigh.

Luna, at high, gains 5.4 points on its predecessor and cuts cost per task by 58%. The absolute score is not on the page.

On Agents’ Last Exam, Sol at max scores 56.4%, above Opus 5’s best on that evaluation and 60% cheaper per task. In the Astra launch post I wrote down 59.3% for the big 6. This page does not reprint that cell. Sol comes close to that earlier number.

On DeepSWE 1.1, Sol at max scores 68.8%. Fable 5 at xhigh, 69.9%: within a point, and about 80% cheaper per task. They use Fable 5 because, they say, they didn’t have the 5.1 score. Luna at max, 66.6%, in the range of Opus 5 and Fable 5 when those run at medium: 93% cheaper than Opus and 96% cheaper than Fable. The effort is not the same. The word “comparable” lives on that mismatch.

On OSWorld 2.0, partial reward on the offline set (the v2026.08.08 release), Sol at xhigh ties Opus 5 at medium: 60.5% against 60.3%, at about 80% less per task. Luna at max beats 5.6 Sol at medium for a tenth of the cost. Astra is still, by their account, the best at using a computer. The 72.6% in the Astra post is a different measurement.

On their internal factuality check — de-identified real conversations where someone had already flagged an error — Sol makes about half as many mistakes as its predecessor and approaches Astra. Luna, at higher effort, matches 5.6 Sol at about a hundredth of the cost. This is not an ordinary day: these are the threads that had already failed. They say answer length barely moves the score.

On alignment, they improve on 5.6, including on made-up claims about what they did while coding. The system card they link is Astra’s. The tests are hard situations, not the failure rate of a normal day.

What someone who didn’t write the announcement says

Artificial Analysis, the same day, at max effort. The intelligence index stays level with 5.6. What moves is the bill.

Sol costs $1.06 per task on that index, against $1.99 for 5.6 Sol. Luna, $0.07 against $0.18. Both write a little more: 31,000 output tokens against 29,000 for Sol, and 51,000 against 41,000 for Luna. The discount absorbs the extra.

On the coding-agent index, measured inside the Codex harness, Sol scores 57, two points above 5.6 Sol. Terminal-Bench 4.0, in that cell, goes from 37% to 43%. SWE-Atlas-QnA, from 54% to 58%. $2.99 per task, about half, and on the cost frontier. Luna scores 41, two points below 5.6 Luna. SWE-Atlas falls from 49% to 44%. DeepSWE, from 66% to 64%. About 60% cheaper, and a bit worse at code.

On hallucinations, AA-Omniscience, both drop: Sol from 92% to 60%, Luna from 93% to 77%. Sol gets there by staying quiet more often. It attempts 83% of the questions; 5.6 attempted 99%. Wrong answers fall by about a quarter, and accuracy falls too, from 59% to 54%. Luna holds accuracy, 44% against 43%, and attempts fewer. Sol’s index goes from 22 to 27. Luna’s, from −10 to 1.

There are regressions. On GDPval-AA, work across 44 occupations, Sol loses about 100 Elo and Luna about 75. On AA-Briefcase, multi-week projects, Luna loses about 45 and Sol stays level. They looked at hundreds of deliverables: the drop comes from poorer presentation and from pieces that leave out what the rubric asked for.

OpenAI sells a similar shortness as a virtue. They bring Astra’s style down: less jargon, less low-value detail, slightly shorter answers. It can read more clearly. It can also be the deliverable missing a section. Both readings fit on the same page.

In the intelligence-index breakdown, Sol’s Terminal-Bench 4.0 comes out 44% against 40%, not the 43% against 37% from the coding index. Same test, another cell. Both rise a little. I don’t put them next to Astra’s 57.7% from September: another harness, another effort.

The cache, which is where an agent shows up

Alongside the price, OpenAI says GPT-6 caching hits more often by default. A cached read is still 10% of the input price. What’s new, for anyone chaining tools, is that changing effort mid-conversation no longer breaks the cache, and turning tools on or off doesn’t either. Explicit breakpoints let you choose which prefix gets stored.

GitHub, by their account, has cut by more than 50% the share of the prompt that has to be processed fresh, across billions of requests and over the past several months. That is not a figure from tonight. It’s the argument they use to say a long agent run costs less than a fresh-token price implies.

What it means

I stay where I am. Yesterday I wrote that 4.7 matters to me if the Moodle plugin comes out cleaner and the week’s bill doesn’t move. Sol’s input is already at that $2. The output is still more expensive, and the place it lives is Codex and ChatGPT Work. Astra was not coming down into Cursor’s model menu. This announcement doesn’t name Cursor either.

Today’s map, read next to this morning’s: Opus 5.5 put Fable-class work at $4 and $20. Sol, the middle GPT-6, is $2 and $10. Both labs cheapened that middle model and left the expensive one on top. The outside index says Sol’s brain barely moved relative to 5.6. The bill did.

Luna is the other bet: volume. On Artificial Analysis’s coding index it drops two points and writes more. On OpenAI’s bench, at high, it improves and costs a lot less. For a lot of cheap work, I’d look at Luna. For the repo, I stay where I was.


Sources: OpenAI — Introducing GPT-6 Sol and Luna, post, Tibo, API pricing, GPT-6 Sol, GPT-6 Luna, Artificial Analysis.

CompartirShare