SpaceXAI has launched Grok 4.7, its newest large language model aimed at developers and knowledge workers. The company says the model is faster and cheaper than its rivals. Grok 4.7 is priced at $2 per million input tokens and $6 per million output tokens, and a fast variant doubles the output speed at twice the price.
The model is available today in Cursor and Grok Build, along with the Grok API, third-party coding harnesses, and model routers and cloud platforms. SpaceXAI says the release is its most capable model for coding and knowledge work yet.
The New Base Model
Grok 4.7 uses a larger base model than Grok 4.6. It was trained with a longer reinforcement learning run on a harder mix of tasks, weighted toward problems that take many hours to complete. That training focused on long-running coding jobs, the kind that stress a model’s ability to hold context over time.
The model is also better at verifying its own work. SpaceXAI says it checks its answers more carefully than earlier versions.
Safety Stacks and Cybersecurity Partnerships
Grok 4.7 was built with an entirely new safeguard stack. SpaceXAI says it is the strongest model the company has tested on refusals and jailbreak resistance.
In dual-use domains like cybersecurity and biological work, the model leads on both utility for benign tasks and safe refusal on dangerous ones. On LatchBio’s biosafety benchmark, Grok 4.7 tops the chart at 62.4%.
The model shows the highest safety on HackerBench v0.3, SpaceXAI’s benchmark for risky and malicious cyber tasks. It allows only 3.3% of risky dual-use prompts through while rarely blocking legitimate security work. SpaceXAI has also started giving select cybersecurity partners invite-only access to Grok 4.7’s red-team capabilities for defense research.
Pricing and Performance Benchmarks
Grok 4.7 is priced the same as Grok 4.6, at the same speed. SpaceXAI calls it highly competitive in its class. On CursorBench 4.0, which stresses longer-running coding tasks, the model is at the frontier in price-performance.
The company also cites GDPval and AA Briefcase as evidence of improved general knowledge work. Those benchmarks ask AI to perform tasks typically done by professionals such as lawyers, nurses, and financial analysts. Grok 4.7 performs comparably to other frontier models on both.
The Shift From Grok 4.6
The move from Grok 4.6 to Grok 4.7 is not just a new version number. The model is larger, the training run is longer, and the focus is on harder tasks.
The sequence of changes, in order:
- A larger base model than Grok 4.6.
- Longer reinforcement learning on a harder mix of tasks.
- Training to verify its own work more carefully.
- Native understanding of the Grok Bot harness for conversational tasks.
- New safeguard stack with stronger refusals and jailbreak resistance.
- Better performance on GDPval and AA Briefcase benchmarks.
- Top spot on LatchBio’s biosafety benchmark at 62.4%.
- Strongest showing on HackerBench v0.3 at 3.3% risky prompt allowance.
- Invite-only access for select cybersecurity partners.
The Cybersecurity Partnership Detail
The invite-only cybersecurity partnerships are worth noting. SpaceXAI is giving select partners access to Grok 4.7’s red-team capabilities for defense research. That is a practical application of the model’s improved safety features.
Red-team capabilities are designed to test systems for weaknesses. By sharing those tools with trusted partners, SpaceXAI is putting its model through real-world adversarial testing.
How the Model Works Today
Grok 4.7 is live now across Cursor, Grok Build, the API, and partner platforms. The pricing is simple: $2 per million input tokens, $6 per million output tokens, with a fast variant doubling the output speed at twice the price.
The model is designed for coding and knowledge work. It is meant to handle long-running tasks without losing context, and it is calibrated to refuse requests that cross the line into dangerous territory.
Why This Matters
The release lands with a clear value proposition. Twice as fast, at half the price of comparable models. That framing is central to how SpaceXAI is pitching the product.
The benchmarks back up the claims. CursorBench 4.0 shows price-performance leadership. GDPval and AA Briefcase show improved professional task performance. LatchBio and HackerBench show the safety stack holds up.
| Benchmark | Focus | Result |
|---|---|---|
| CursorBench 4.0 | Long-running coding tasks | Frontier in price-performance |
| GDPval | Professional tasks (lawyers, nurses, analysts) | Improved on Grok 4.6 |
| AA Briefcase | Professional tasks (lawyers, nurses, analysts) | Improved on Grok 4.6 |
| LatchBio | Biosafety | 62.4% top spot |
| HackerBench v0.3 | Cybersecurity | 3.3% risky prompt allowance |
The Bottom Line
Grok 4.7 is a technical upgrade with a clear business pitch. It is faster, cheaper, and safer than its predecessors. The benchmarks are public, the pricing is transparent, and the model is shipping now.
The question for developers and knowledge workers is whether the improvements matter for their actual work. Longer reinforcement training on harder tasks means the model should hold context better on long coding sessions. The safety stack adds guardrails for sensitive domains.
The partnership program with cybersecurity firms is a notable move. It gives trusted partners access to Grok 4.7’s red-team capabilities for defense research. The model is available now, the benchmarks are published, and the price is set.
Source material: “Grok 4.7,” x.ai.
Get the Notebook.
The day's best stories and every fresh verdict, in plain English, in your inbox by seven. One email a day, no more.

