Grok 4.7 is out: same price, the longer task
ContenidoContents
On 16 August I wrote that Musk had 4.7 three or four weeks out, and that we’d see: Elon promises dates the way other people order coffee. Today, 21 September, SpaceXAI shipped it. This is not the verdict of a week inside Moodle. It’s the spec sheet, read slowly.
What it is, no fluff
Grok 4.7 is, by their account, their most capable model for coding and knowledge work. The base is new and larger than 4.6. The reinforcement-learning run was longer, on a harder mix, weighted toward problems measured in hours. They say it checks its own work more carefully, handles long context better, and natively understands the Grok Bot harness.
The concrete part, from the docs:
- ID:
grok-4.7. - Context: 500,000 tokens. The same ceiling as 4.6.
- Knowledge cutoff: May 2026.
- Input: text and image. Output: text, with no published cap.
- Effort:
low,medium,high(the default) orxhigh. - API tools: function calling, web search, X search and code execution.
Today it runs in Grok Build, in Cursor (every plan), on the API, and on OpenRouter, Vercel and Cloudflare. There is a fast variant, twice the output speed, only in Cursor and Grok Build. It is not on the public API, and it is not in Grok Build’s free tier.
The price, which is 4.6’s price
The announcement headline says “twice as fast, at half the price of comparable models.” Against 4.6, no: same price and, they say, same speed. The half is against the row next door.
Per million tokens, with the prompt under 200,000, from the price list:
| Model | Input | Output |
|---|---|---|
| Grok 4.7 | $2 | $6 |
| Grok 4.6 | $2 | $6 |
| GPT-5.6 Sol | $4 | $20 |
| Fable 5.1 | $10 | $50 |
Cached input on 4.7 is $0.50. In an agent looping over the same repo, that matters more than the headline.
The 200,000 threshold does not surcharge only the overflow. Once the prompt crosses it, every token in that request bills at double: $4 input and $12 output. Same as 4.6. A long Grok Build thread can cross it without a warning.
The fast variant, under 200,000, is exactly double: $4 and $12. Above that, the table says $6 input and $18 output. That is one and a half times the $4 / $12 long-context rate, not the double suggested by the docs’ phrase (“twice the standard token rates”). The table wins.
The United States endpoint (us.api.x.ai) charges 10% more and, for now, only serves 4.7 and 4.6.
And a detail that gets expensive if you skip it: they recommend setting prompt_cache_key (or the x-grok-conv-id header on Chat Completions). Without it, the conversation hops servers and you pay full input even when the prefix has not changed.
The table, with the effort dial set
The percentages are SpaceXAI’s. The 4.7 column is at xHigh. The 4.6 column is at High. Sol and Fable 5.1 are at Max. Not the same dial. The DeepSWE asterisk also marks that 71% as high effort, not xHigh.
| Benchmark | 4.7 | 4.6 | Sol | Fable 5.1 |
|---|---|---|---|---|
| CursorBench 4.0 | 46.3% | 40.4% | 41.7% | 51.8% |
| DeepSWE v1.1 | 71.0%* | 65.2% | 72.7% | 70.0% |
| EEBench | 64.0% | 53.0% | 39.4% | 56.4% |
| AA Briefcase v1.1 | 1,657 | 1,546 | 1,487 | 1,678 |
| Terminal-Bench 4.0 | 38.0% | 20.3% | 37.3% | 57.9% |
| Harvey (legal) | 19.6% | 15.8% | 2.5% | 6.7% |
| HealthBench Professional | 56.7% | 48.5% | 60.5% | 62.1% |
* high effort
How I read it, row by row.
On CursorBench, the long-running coding bench they highlight, it beats 4.6 and Sol and trails Fable. That is where they say they sit at the price-performance frontier. At $2 / $6 against $10 / $50, the argument is not the percentage. It is the percentage per dollar.
On DeepSWE the 71% lands under Sol and above Fable, with the effort asterisk. It does not support “won the row.”
EEBench, electrical engineering, is the widest jump over 4.6: eleven points, and ahead of Fable. That is the row that looks like a new model, not a tune.
AA Briefcase, multi-hour office work, is up on 4.6 and on Sol and a step behind Fable (1,657 against 1,678). Knowledge work, not code. Comparable.
Terminal-Bench 4.0 nearly doubles 4.6, from 20.3% to 38.0%, and ties Sol. Fable goes to 57.9%. Three weeks ago, at the GPT-6 Astra launch, OpenAI put the 6 at 57.7% on that same bench. 4.7 is not in that conversation. It is in Sol’s, at a lower price.
Harvey it wins clearly, and the absolute number is still low: 19.6%. A legal agent that lands one in five is not a row to celebrate. It is a row where everyone else is worse.
HealthBench improves on 4.6 and trails Sol and Fable. Clinical work is not the pitch.
On safety they describe a new stack. They say it is the strongest model they have tested on refusals and on resistance to jailbreaks. On LatchBio’s biosafety benchmark they mark 62.4%. On HackerBench v0.3, theirs for risky dual-use cyber prompts, 3.3% of the risky ones get through and, they say, legitimate security work is rarely blocked. Selected partners have invite-only access to the red-team capabilities. Their figures. The method is not in the post.
What I’ll actually watch
I’ve been coding with Grok all summer. I started out of curiosity and I stayed because it stays in the repo: one thread, the diff, the commit, the Hugo build. 4.6 was already enough for that.
4.7 matters to me if it lasts longer inside a Moodle plugin without inventing the third file, and if the week’s bill does not move. The list price has not moved. What can move is the effort dial: if Grok Build climbs to xHigh because it now can, the invoice is not the one in the announcement.
I’m leaving the fast variant parked. Double the price for double the speed, and only inside Cursor or Grok Build. For the weekday, plain 4.7.
I don’t switch labs over a table. If in a few days the code comes out cleaner on the first pass, I’ll say so. If it’s 4.6 with a new number, I’ll say that too.