Grok 4.7 Review: Cheap at $2/$6, but Is It Worth Switching from GPT-6 or Claude Fable 5.1?
What's new
xAI released Grok 4.7 on September 21, 2026, calling it its most capable model for coding and knowledge work. According to xAI's announcement, it is built on a new, larger base model than Grok 4.6, trained with a longer reinforcement-learning run weighted toward multi-hour tasks, and is better at checking its own work and handling long context.
The headline is the price. Grok 4.7 is served at the same price and speed as Grok 4.6:
| Model | Input ($/M tokens) | Output ($/M tokens) |
|---|---|---|
| Grok 4.7 | $2 | $6 |
| GPT-5.6 Sol | $4 | $20 |
| Claude Fable 5.1 | $10 | $50 |
xAI also offers a fast variant with twice the output speed at twice the price. At launch it is available through the Grok API, Cursor, Grok Build, third-party coding harnesses and model routers.
The benchmarks: xAI's numbers vs independent tests
xAI's own table (Grok 4.7 at xHigh effort vs rivals at max effort):
| Benchmark | Grok 4.7 | Grok 4.6 | GPT-5.6 Sol | Fable 5.1 |
|---|---|---|---|---|
| CursorBench 4.0 (long coding tasks) | 46.3% | 40.4% | 41.7% | 51.8% |
| DeepSWE v1.1 | 71.0%* | 65.2% | 72.7% | 70.0% |
| EEBench (electrical engineering) | 64.0% | 53.0% | 39.4% | 56.4% |
| AA Briefcase v1.1 (multi-hour office work) | 1,657 | 1,546 | 1,487 | 1,678 |
| Terminal-Bench 4.0 | 38.0% | 20.3% | 37.3% | 57.9% |
| Harvey Legal Agent Benchmark | 19.6% | 15.8% | 2.5% | 6.7% |
| HealthBench Professional | 56.7% | 48.5% | 60.5% | 62.1% |
* Grok 4.7 DeepSWE score at high effort.
Two things stand out. First, Grok 4.7 is a genuine upgrade over 4.6 on every line, and it leads on legal agent work and electrical engineering. Second, xAI compared it against GPT-5.6 Sol, not OpenAI's current flagship GPT-6 Astra.
Independent testing is less generous. On the Artificial Analysis Intelligence Index (v4.3.2, ten benchmarks combined), Grok 4.7 scores 46 — mid-pack — while Claude Fable 5.1 and GPT-6 lead with 53 each. The gap is widest in agentic coding: in The Decoder's report of independent Terminal-Bench 4.0 results, Grok 4.7 reaches 26%, versus 60% for GPT-6 Astra and 55% for Claude Fable 5.1 — and even the much cheaper DeepSeek V4.1 Flash edges past it at 27%.
So the fair summary is: near-frontier on knowledge work, clearly behind on long-running agentic coding, priced like a budget model.
Can you try it on AIWITH.CHAT?
Not yet — Grok 4.7 is not on AIWITH.CHAT, and we'd rather say so than pretend. What is there today is every model xAI and the independent testers compared it against:
- Claude Fable 5.1 — the model Grok 4.7 is chasing on CursorBench, Terminal-Bench and office work
- GPT-6 Astra and GPT-5.6 Sol — the GPT-6 flagship that leads Terminal-Bench, plus the model in xAI's own chart
- DeepSeek V4.1 Flash — the budget model that matched Grok 4.7 on independent Terminal-Bench
All of them sit in one dropdown under a single $9.9/month plan, so you can run the same prompt through Fable 5.1, GPT-6 Astra and DeepSeek V4.1 Flash and see which one actually handles your task — without paying per-token API rates.
Compare them side by side on AIWITH.CHAT →
Is it worth switching?
If you're a developer paying API bills for Claude Fable 5.1 or GPT-5.6 Sol: Grok 4.7 is worth a test run on your non-agentic workloads — code review, refactors, document drafting. At $6 vs $50 per million output tokens it is roughly 8x cheaper than Fable 5.1 on output, and on DeepSWE and AA Briefcase it is within a couple of points.
If you rely on long autonomous coding sessions (terminal agents, multi-step repo tasks): no. Independent Terminal-Bench puts it at less than half of GPT-6 Astra and Fable 5.1. Saving on tokens doesn't help if the agent needs three times as many retries.
If you do legal or engineering document work: this is the one area where Grok 4.7 leads outright on xAI's numbers, so it's worth trying — but verify on your own documents, since these are vendor-run benchmarks.
If you're a chat user, not an API buyer: per-token price doesn't reach you. What matters is which model answers your questions best, and a multi-model subscription lets you pick per task instead of betting on one vendor.
Related reading
- Claude Fable 5.1 free access: how to try it without a Claude Pro plan
- GPT-6 Astra free: how to try OpenAI's most agentic model
- GPT-6 Astra vs Claude Fable 5 for coding
- DeepSeek V4.1 Flash: what the KV-cache change means
- AI Model Costs Explained: Why Some Chats Cost 10x More
Sources: xAI — Introducing Grok 4.7 · The Decoder — xAI launches Grok 4.7 at bargain prices, but benchmarks reveal a wide gap to Claude and GPT-6