Grok 4.7 targets code and keeps the same per-token price as its predecessor
Grok 4.7 was launched by SpaceXAI on September 21, with a focus on coding and knowledge work. The company promises half the operational cost.

Grok 4.7 was launched by SpaceXAI on September 21, with a focus on coding and knowledge work. The company promises half the operational cost of rival systems. The per-token price remains the same as Grok 4.6's.
Prices stand at US$ 2 per million input tokens. For output, the charge reaches US$ 6 per million.
There is also an accelerated variant. However, it doubles generation speed and costs twice the standard price.
Grok 4.7 bets on reinforcement training and self-verification
The architecture starts from an expanded base. On top of it, however, the company ran extended reinforcement learning training.
Furthermore, the focus of this computing effort stands out. The teams concentrated their efforts on complex, multi-hour problems.
In parallel, the model gained native self-verification mechanisms. Context handling also grew.
According to the company, the model works longer on difficult tasks and checks its own work more carefully. The company also cites the best-calibrated safeguards it has ever delivered.
However, engineers also integrated direct compatibility with the Grok Bot harness. As a result, dialogue and unstructured knowledge processing received adjustments.
The coding numbers show progress with a clear ceiling
However, on CursorBench 4.0, which measures long coding tasks, the model scored 46.3% success. The average cost was US$ 11.95 per completed task.
This result surpasses two direct competitors. However, Grok 4.6 scored 40.4% and GPT 5.6 Sol scored 41.7%. Still, Fable 5.1 reached 51.8%.
On DeepSWE v1.1, the scenario changes. The model reached 71% in a high-effort configuration, ahead of Grok 4.6's 65.2% and Fable 5.1's 70%.
The most visible jump appeared on Terminal Bench 4.0. The score rose from 20.3% to 38% between generations.
Grok 4.7 stands out in specific professional domains
On GDPval, the model registered 1,695 Elo points. It trailed Fable 5.1, at 1,735, and was ahead of Grok 4.6 and GPT 6 Astra.
In electrical engineering, the gap was wide. EEBench showed 64.0% versus 39.4% for GPT 5.6 Sol.
On the Harvey Legal Agent Benchmark, the result reached 19.6%. The others fell well below, with Grok 4.6 at 15.8% and Fable 5.1 at 6.7%.
In the clinical domain, performance landed in the middle. HealthBench Professional recorded 56.7%, above its predecessor and below GPT 5.6 Sol and Fable 5.1.
Notice the pattern in these numbers. No model leads every category, and the choice depends on your domain.
The safety layer was rebuilt
The company redid the defensive filtering architecture. The stated goal addresses dual-use risks while preserving legitimate technical research.
On LatchBio's biological safety benchmark, the model leads with 62.4%. This score combines usefulness on benign queries with refusal on dangerous topics.
On HackerBench v0.3, which covers destructive cyber operations, the rate of improper completions stood at 3.3%. At the same time, legitimate security workflows remain allowed.
Selected cybersecurity partners received restricted access. They will assess red-team capabilities for defensive research.
Where to access the model today
General distribution has already begun across commercial channels. Access is available through the Grok API, third-party model routers, and managed cloud environments.
Cursor and Grok Build have granted immediate access. Terminal installers have also arrived through standard shell environments.
What to evaluate before switching models
Start with the cost per completed task. The per-token price matters little when the model needs many attempts.
Then, test with your own repository. Public benchmarks indicate direction, but your code decides the choice.
Also consider your team's dominant task. Those who live in the terminal should watch the Terminal Bench jump closely.
Also, measure real-world latency. The fast variant doubles the price, and that only pays off in time-sensitive workflows.
Finally, keep the abstraction layer. Switching providers remains the cheapest decision when price or policy changes.
Follow our profile on Instagram!
Translated from the Brazilian Portuguese original · Read the original
Windows Zenith targets the dev who currently chooses Linux
Windows Zenith emerges as Microsoft's bet on a system built for development and local AI. The proposal circulated on September 21.






