Artificial Analysis via THE DECODER SpaceXAI's Grok 4.7 Is Here: Bigger Base Model, Longer RL, and a Mid-Pack Reality Check
SpaceXAI released Grok 4.7 on September 21 — a larger base model with longer reinforcement learning aimed at coding and agentic work, priced at $2/$6 per million tokens. Vendor benchmarks show gains over Grok 4.6, but independent tests put it well behind Claude Fable 5.1 and GPT-6.
SpaceXAI released Grok 4.7 on September 21, the first frontier model to ship under the combined SpaceXAI banner since xAI’s merger into SpaceX. The company describes it as a larger base model than Grok 4.6, trained with longer reinforcement learning on harder, multi-hour coding and agentic tasks, and designed to better verify its own output. Pricing is unchanged from 4.6: $2 per million input tokens and $6 per million output tokens, with the rate doubling above 200K tokens of context, per aireleasetracker. OpenRouter lists a cheaper route at $1.60/$4.80.
What SpaceXAI claims
The vendor-reported numbers show a clear step up from Grok 4.6. On CursorBench 4.0, Grok 4.7 scores 46.3% versus 40.4% for its predecessor. On DeepSWE v1.1 it reaches 71.0% at high reasoning effort against 65.2%, and on OSWorld-verified it posts 70.3% versus 61.2%, according to sqmagazine’s breakdown. The model ships with selectable reasoning effort — low, medium, high, and xhigh — and reports put the context window at 500K tokens. It is available through Cursor on all tiers, the Grok API, Grok Build (free at x.ai/build), OpenRouter, and third-party harnesses.
SpaceXAI also says Grok 4.7 carries a new safeguard stack built for agentic misuse, citing 62.4% on a LatchBio biosafety test and a 3.3% risky pass-through rate on its own HackerBench v0.3 — that last one being the vendor’s own benchmark, so treat it accordingly. Elon Musk’s pre-launch claim of 2.1 trillion parameters remains unconfirmed by the model card.
What independent tests show
The independent picture is less flattering. On the Artificial Analysis Intelligence Index (v4.3.2), which blends ten benchmarks, Grok 4.7 scores 46 — mid-pack, against 53 each for Claude Fable 5.1 and GPT-6. The gap widens on agentic coding: Terminal-Bench 4.0 gives it 26%, versus 60% for GPT-6 Astra, 55% for Claude Fable 5.1, and even 27% for the cheaper DeepSeek V4.1 Flash, as THE DECODER reports.
That pricing context matters. At $2/$6, Grok 4.7 costs roughly what Chinese labs charge — well under Western frontier rates — but the independent benchmarks suggest it also performs closer to the Chinese tier on agentic work. For developers, the practical question is whether the Cursor distribution and aggressive pricing make it a useful secondary model for high-volume coding tasks, rather than a frontier replacement. The Terminal-Bench numbers say: keep Claude or GPT-6 in the loop for the hard stuff.
The Cursor angle deserves emphasis. Shipping on all Cursor tiers on day one puts Grok 4.7 directly in front of the agentic-coding audience SpaceXAI is targeting, and the free Grok Build tier at x.ai/build gives curious developers a zero-cost way to test the xhigh reasoning effort setting against their own repos. If you do try it, benchmark it on your own multi-hour tasks rather than trusting either the vendor’s CursorBench numbers or the independent Terminal-Bench scores — the two disagree by enough that your workload is the only tiebreaker that matters.