SPIRITT logoSPIRITTSpaceXAIGrok 4.7

Grok 4.7 is built for harder work

A larger model. Longer training runs. More careful verification. SpaceXAI's Grok 4.7 raises the bar for coding and professional work while keeping standard API pricing at $2/M input and $6/M output. Explore the release, then bring your next project to a SPIRITT Workspace.

Grok 4.7 on X

See what people are building with Grok 4.7, from 3D worlds and video edits to apps and everyday workflows.

A larger model, built to stay with the task

Grok 4.7 moves to a new, larger base model and a longer reinforcement-learning run on harder, multi-hour tasks. The focus is not just producing an answer, but checking the work and managing a longer trajectory.

More capability at the same base rate

SpaceXAI reports improvements over Grok 4.6 across every benchmark in its release table, spanning software engineering, terminal work, legal tasks, clinical reasoning and electrical engineering. Grok 4.7 is stronger, but Fable 5.1 still leads several of the displayed coding and professional-work comparisons.

The API supports a 500,000-token context window and low, medium, high and xhigh reasoning effort. Standard pricing below 200K input tokens is $2/M input, $0.50/M cached input and $6/M output. At 200K input tokens or more, the entire request moves to $4/$1/$12. Default API effort is high; the launch comparison mostly uses xhigh.

Multi-hour codingProfessional knowledge work500K contextSelf-verificationLow to xhigh reasoningStronger safeguards

CursorBench 4.0

Near frontier
Fable 5.1 max
51.8%
Grok 4.7 xhigh
46.3%
GPT-5.6 Sol max
41.7%
Grok 4.6 high
40.4%

SpaceXAI release table, September 21, 2026. Vendor-reported comparison; model efforts differ as labeled. Grok improves 5.9 percentage points over 4.6; Fable leads on raw score. The release positions Grok on the price-performance frontier.

The full release-table comparison

Seven task benchmarks, a separate professional-work chart, and the published base token rates. Vendor-reported results stay labeled; price, reasoning effort and benchmark version are part of the comparison.

Official software engineering
Near frontier

DeepSWE v1.1

GPT-5.6 Sol max
72.7%
Grok 4.7 high
71%
Fable 5.1 max
70%
Grok 4.6 high
65.2%

SpaceXAI release table, September 21, 2026. Vendor-reported comparison; model efforts differ as labeled. Important: Grok 4.7 uses high effort here, not xhigh.

Official electrical engineering
Winner

EEBench

Grok 4.7 xhigh
64%
Fable 5.1 max
56.4%
Grok 4.6 high
53%
GPT-5.6 Sol max
39.4%

SpaceXAI release table, September 21, 2026. Vendor-reported comparison; model efforts differ as labeled. Grok 4.7 has the highest score in this table.

Official multi-hour office work
Near frontier

AA Briefcase v1.1

Fable 5.1 max
1678
Grok 4.7 xhigh
1657
Grok 4.6 high
1546
GPT-5.6 Sol max
1487

SpaceXAI release table, September 21, 2026. Vendor-reported comparison; model efforts differ as labeled. Elo scores. These are the values in the vendor table, not an independently rerun evaluation.

Official multi-hour terminal work

Terminal-Bench 4.0

Fable 5.1 max
57.9%
Grok 4.7 xhigh
38%
GPT-5.6 Sol max
37.3%
Grok 4.6 high
20.3%

SpaceXAI release table, September 21, 2026. Vendor-reported comparison; model efforts differ as labeled. Grok gains 17.7 points over 4.6. Fable 5.1 remains clearly ahead; do not mix these results with older Terminal-Bench versions.

Official legal work
Winner

Harvey Legal Agent Benchmark

Grok 4.7 xhigh
19.6%
Grok 4.6 high
15.8%
Fable 5.1 max
6.7%
GPT-5.6 Sol max
2.5%

SpaceXAI release table, September 21, 2026. Vendor-reported comparison; model efforts differ as labeled. Leading this displayed benchmark is not a substitute for professional review of legal work.

Official clinical reasoning

HealthBench Professional

Fable 5.1 max
62.1%
GPT-5.6 Sol max
60.5%
Grok 4.7 xhigh
56.7%
Grok 4.6 high
48.5%

SpaceXAI release table, September 21, 2026. Vendor-reported comparison; model efforts differ as labeled. Grok improves over 4.6 but trails Fable and Sol. Benchmark performance does not establish suitability for autonomous clinical decisions.

Official professional knowledge work
Near frontier

GDPval

Fable 5.1 max
1735
Grok 4.7 xhigh
1695
Grok 4.6 high
1605
GPT-6 Astra max
1542

SpaceXAI interactive launch chart, September 21, 2026. Elo scores. This chart compares GPT-6 Astra, unlike the release table above, which compares GPT-5.6 Sol.

Published base API economics

Input price per million tokens

Grok 4.7
$2.00
Grok 4.6
$2.00
GPT-5.6 Sol
$4.00
Fable 5.1
$10.00

SpaceXAI release-table prices. Lower is better. Grok 4.7 retains the same base rate as 4.6. Long-context, cached-input and priority pricing are separate.

Published base API economics

Output price per million tokens

Grok 4.7
$6.00
Grok 4.6
$6.00
GPT-5.6 Sol
$20.00
Fable 5.1
$50.00

SpaceXAI release-table prices. Lower is better. Token rate is not total task cost: reasoning tokens, context length and the number of attempts still matter.

Source: SpaceXAI's September 21, 2026 Grok 4.7 launch report and current xAI API documentation. All charts above are attributed vendor-reported results, not SPIRITT reruns. The release table uses GPT-5.6 Sol; GDPval uses GPT-6 Astra. Most Grok 4.7 scores use xhigh effort, while DeepSWE uses high. Standard prices shown exclude long-context and priority premiums. Full sources are linked below.

How It Works

Give the work a proper home: files, tools, a browser and a concrete finish line.

01

Open a real workspace

Bring your repository, files, research or idea into SPIRITT. Your agent has a browser, terminal, integrations and memory, so the task can move beyond a conversation.

A cloud workspace with connected browser, files and terminal
02

Match Grok 4.7 to the task

Use the release evidence to judge the fit: sustained coding, document work, verification loops and 500K context. Choose the model and access route available in your workspace before committing a real workflow.

Grok model identity in the established pastel workspace illustration
03

Build, check, and finish

Set a clear outcome and ask your agent to test its work. Keep the source files, results and decisions together, with human review wherever the stakes demand it.

Build and automation tools in a connected SPIRITT workspace

Put your next project to work

Give your AI agent a workspace with the tools, files and memory to reach a real finish line.

Questions

Frequently asked questions

01What changed from Grok 4.6?+
SpaceXAI describes a new, larger base model and a longer reinforcement-learning run on harder tasks that can take many hours. Grok 4.7 is trained to verify its work more carefully, manage long context and understand the Grok Bot harness. The standard base token rates remain unchanged.
02Does Grok 4.7 lead every benchmark?+
No. It leads the displayed EEBench and Harvey Legal comparisons, but Fable 5.1 scores higher on CursorBench 4.0, Terminal-Bench 4.0, HealthBench Professional, AA Briefcase and GDPval. GPT-5.6 Sol leads the displayed DeepSWE comparison. The useful question is which model fits your actual workload and budget.
03How much does the standard Grok 4.7 API cost?+
For prompts below 200,000 tokens, published prices are $2 per million input tokens, $0.50 per million cached input tokens and $6 per million output tokens. Prompts of 200,000 tokens or more apply $4/$1/$12 rates to all tokens in the request. These are xAI API rates, not a SPIRITT subscription price.
04What is Grok 4.7 Fast?+
The docs describe the same model served on faster infrastructure, currently limited to Cursor and Grok Build rather than the public xAI API. Below 200K input tokens its listed rates are $4/$1/$12 per million input/cached/output tokens; above 200K the table lists $6/$1.50/$18. The Fast docs do not clearly specify the exact 200K boundary. Public API Priority Processing is a separate feature.
05How large is the context window, and can reasoning be disabled?+
The documented context window is 500,000 tokens. Reasoning effort can be low, medium, high or xhigh; the API defaults to high. Reasoning cannot be disabled. An API default does not tell you which effort was used for a benchmark.
06What do the safety claims mean?+
SpaceXAI reports a new safeguard stack, a 62.4% LatchBio biosafety score and a 3.3% risky-prompt allowance rate on its HackerBench v0.3 evaluation. These are vendor claims about specific tests, not a universal safety guarantee. Sensitive legal, medical and security work still needs qualified oversight.
07How should I check Grok 4.7 access in SPIRITT?+
Start with a SPIRITT Workspace and confirm the models and access routes available there. This page documents SpaceXAI's verified Grok 4.7 release; it does not guarantee that Grok 4.7 is already selectable in every workspace. The linked API prices describe xAI's service, not a SPIRITT subscription.
Buy from builders who use what they sellBuilt usingSPIRITT