DeepSWE v1.1
SpaceXAI release table, September 21, 2026. Vendor-reported comparison; model efforts differ as labeled. Important: Grok 4.7 uses high effort here, not xhigh.
A larger model. Longer training runs. More careful verification. SpaceXAI's Grok 4.7 raises the bar for coding and professional work while keeping standard API pricing at $2/M input and $6/M output. Explore the release, then bring your next project to a SPIRITT Workspace.
See what people are building with Grok 4.7, from 3D worlds and video edits to apps and everyday workflows.
Grok 4.7 is here. It's a notable improvement over Grok 4.6 at the same price and speed.
— SpaceXAI (@SpaceXAI) 2026-09-21T16:17:53Z
Compare Grok 4.7 (first) and 4.6 (second) building an open world city game.
— SpaceXAI (@SpaceXAI) 2026-09-21T16:17:54Z
grok 4.7 is here, and its our best model so far! try it out in cursor, grok build, api or anywhere you get your tokens! curious to hear what you think here's grok 4.6 vs 4.7 building age of empires ii https://t.co/4gWkFev4zO
— eric zakariasson (@ericzakariasson) 2026-09-21T16:20:09Z
i was able to test Grok 4.7 a few days early (very thankful) it built the whole golden gate bridge in Blender without downloading any meshes/textures online it took 17 minutes, here was the prompt: "use the blender MCP to make a detailed model of the golden gate bridge. do not pull in assets from the web." then it made this video
— Anthony Kroeger (@kr0der) 2026-09-21T16:24:55Z
Grok 4.7 is here! And it's good at Video Editing???? watch here: https://t.co/hMnInsOck5 https://t.co/EqLlMkHm0E
— vogel (@ryanvogel) 2026-09-21T16:19:47Z
Grok 4.7 created this 4D video visualizer in 3 prompts! https://t.co/0sNIQ2myiw
— vogel (@ryanvogel) 2026-09-21T16:32:49Z
Trying the new grok on the node.js codebase. @grok 4.7 improves node.js byte-mode pipe() performance by 77%! https://t.co/heNQNwv9AS https://t.co/ihxFj6Z4HX
— Yagiz Nizipli (@yagiznizipli) 2026-09-21T17:14:36Z
Congratulations to @SpaceXAI on the @grok 4.7 drop, we've all been waiting! It's very fast, and the intelligence is quite good. Running multiple cloud agents with it now Spin it up in cloud agents with `railway code --grok` https://t.co/AngHY6vDJl
— Railway (@Railway) 2026-09-21T16:34:12Z
Well well well, Grok 4.7 can do design. Just a single lava lamp test is a good one; when tested with each model, Astra and Fable 5.1 are so far best, but Grok 4.7 is pretty close! https://t.co/QmSgAO4AWY
— Jaanus L. (@JaanusBuilds) 2026-09-21T16:18:00Z
GROK 4.7 IS ACTUALLY COMPETING WITH GPT-6 ASTRA. I gave GPT-6 Astra, Grok 4.7, Kimi K3 and Fable 5.1 the same prompt to build a flight simulator Astra was still #1 overall, but Grok 4.7 was surprisingly close Kimi K3 and Fable 5.1 were basically a draw and both produced a much smoother result, while Grok 4.7 was right up there with Astra in terms of overall quality. overall: Astra > Grok ≈ Kimi ≈ Fable GPT-6 Astra finally has some serious competition.
— aditya (@adxtyahq) 2026-09-21T17:16:40Z
Grok 4.7 is insane in writing front-design I just told Cursor with Grok 4.7 to clone my favorite websites. This is the craft it delivered👇 https://t.co/q1GLcONxgO
— Braxxxx (@Braxxxx_Li) 2026-09-21T17:11:17Z
First test of Grok 4.7 , raw imaginability Prompt : build paintings using three js only which represents different emotions of humans . HTML,CSS and JS site https://t.co/w7QHV5ghvk
— Git Goats (@gitgoats) 2026-09-21T16:18:27Z
Grok 4.7 is a workhorse and easily replaces many models including Opus for a lot of work. Why overpay for the same intelligence when Grok is still $2/$6 vs $5/$25. I use pstack (poteto-mode) to help me fan out multiple agents for research, code review, testing, and implementation. Be sure to tell the model to update your pstack model config to use the latest models. Just prompt the agent and it’ll do it for you.
— Ray Fernando (@RayFernando1337) 2026-09-21T17:14:09Z
I’ve been testing Grok 4.7… it’s a really great daily driver, and a huge step up over 4.6. Definitely worth trying!
— Matt Shumer (@mattshumer_) 2026-09-21T17:15:05Z
Grok 4.7 will be my default model for a lot of tasks > it's fast > it's consistent > it's near-frontier > it's concise Grok 4.6 has been one of the most reliable models I've ever used and Grok 4.7 is just a all-round better version of that https://t.co/8G2U68VDBB
— David Ondrej (@DavidOndrej1) 2026-09-21T16:48:50Z
excited to release Grok 4.7! there's a tension between making a model fast & having it check its own work. we've tried to strike a good balance, with more to do for future releases :) https://t.co/fKiZpSRyOT
— Niklas Muennighoff (@Muennighoff) 2026-09-21T16:56:56Z
Grok 4.7 by @SpaceXAI and @elonmusk is now in the Agent Arena! Your votes drive the @arena leaderboards, head over and bring your toughest prompts. In Agent Arena, we measure models on millions of real-world, long-horizon agentic tasks from a global community of users. Models can access web search, filesystem, and terminal tools to complete complex workflows. The leaderboard measures model performance on outcomes relative to the average model using a causal tracing methodology. In addition to Agent Arena, Grok 4.7 is in Battle Mode for: Text, Vision, Code, and Document.
— Arena.ai (@arena) 2026-09-21T16:37:46Z
Grok 4.7 vs SWE 2 tested both models with same prompt at highest reasoning available and results came out really different > grok 4.7 took 32 minutes and cost $8.14 to make this > swe 2 took 15 minutes and cost $0 to make this really surprised by how good a free model is and no idea what grok was doing here which one did better?
— J A Z I I (@notjazii) 2026-09-21T16:34:38Z
Grok 4.7 scores 46 on the Artificial Analysis Intelligence Index to bring SpaceXAI into the top 4 AI labs. Coding Agent Index performance has also improved, overtaking GPT-5.6 Sol Grok 4.7 scores +2 points over Grok 4.6 on the Intelligence Index, with strong performance on agentic knowledge work tasks. We evaluated the new model at xhigh reasoning effort. Congratulations to @SpaceXAI and @ElonMusk on the release! Key takeaways: ➤ Grok 4.7 joins the frontier of agentic knowledge work: Grok 4.7 gains +111 Elo over Grok 4.6 (high) on AA-Briefcase, our private benchmark for long-horizon agentic knowledge work, scoring 1657 Elo and placing it alongside Claude Opus 5 and Claude Fable 5.1 at the frontier. On GDPval-AA, it scores 1695 Elo, +90 ahead of Grok 4.6 (high). ➤ A leap in coding agent performance: Grok 4.7 (xhigh) with Grok Build scores 56 on the Artificial Analysis Coding Agent Index, up +9 points from Grok 4.6 (xhigh). Among models in their native harnesses, Grok 4.7 + Grok Build now ranks 4th, behind only Claude Fable 5.1, GPT-6 Astra, and Claude Opus 5. ➤ Incremental performance changes elsewhere: Outside of agentic knowledge work, Grok 4.7 broadly matches Grok 4.6 (high) on the other Intelligence Index tasks. It improves on Terminal-Bench 4.0 (+4.5 percentage points) and GDP.pdf (+3.0 p.p.), with regressions on AA-LCR (-3.7 p.p.) and AutomationBench-AA (-1.1 p.p.). ➤ High token use across tasks: Grok 4.7's gains come with higher token usage. Grok 4.7 (xhigh) uses approximately 81k output tokens per Intelligence Index task, compared with 36k for Grok 4.6 (high) and 27k for GPT-6 Astra (max) - 125% and 196% more, respectively. Other model details: ➤ Context window of 500k tokens, unchanged from Grok 4.6 ➤ Pricing of $2/$6 per 1M input/output tokens with cache hits discounted to $0.50 per 1M tokens, matching Grok 4.6 ➤ Configurable reasoning effort spans low to xhigh. Our evaluation uses xhigh.
— Artificial Analysis (@ArtificialAnlys) 2026-09-21T16:38:05Z
Grok Build with Grok 4.7 (xhigh) scores 56 on the Artificial Analysis Coding Agent Index, up from 47 with Grok 4.6 (xhigh). It improves across all three components: DeepSWE v1.1 rises from 65% to 73%, Terminal-Bench 4.0 from 18% to 33%, and SWE-Atlas-QnA from 58% to 63%. These results evaluate Grok with Grok Build, its first party coding agent. They are separate from the Intelligence Index results, which standardize the evaluation harness used across models.
— Artificial Analysis (@ArtificialAnlys) 2026-09-21T16:38:07Z
Our performance measurements put Grok 4.7's answer output speed at approximately 188 tokens/second for long prompts. Grok 4.7 averaged approximately 7.1 minutes per Intelligence Index task.
— Artificial Analysis (@ArtificialAnlys) 2026-09-21T16:38:09Z
Grok 4.7 (xhigh) has a lower AA-Omniscience Hallucination Rate than Grok 4.6 (high): 29% versus 34%.Accuracy is broadly unchanged at 47% versus 48%, and overall AA-Omniscience Index improves from 30 to 32.
— Artificial Analysis (@ArtificialAnlys) 2026-09-21T16:38:11Z
Grok 4.7 moves to a new, larger base model and a longer reinforcement-learning run on harder, multi-hour tasks. The focus is not just producing an answer, but checking the work and managing a longer trajectory.
SpaceXAI reports improvements over Grok 4.6 across every benchmark in its release table, spanning software engineering, terminal work, legal tasks, clinical reasoning and electrical engineering. Grok 4.7 is stronger, but Fable 5.1 still leads several of the displayed coding and professional-work comparisons.
The API supports a 500,000-token context window and low, medium, high and xhigh reasoning effort. Standard pricing below 200K input tokens is $2/M input, $0.50/M cached input and $6/M output. At 200K input tokens or more, the entire request moves to $4/$1/$12. Default API effort is high; the launch comparison mostly uses xhigh.
SpaceXAI release table, September 21, 2026. Vendor-reported comparison; model efforts differ as labeled. Grok improves 5.9 percentage points over 4.6; Fable leads on raw score. The release positions Grok on the price-performance frontier.
Seven task benchmarks, a separate professional-work chart, and the published base token rates. Vendor-reported results stay labeled; price, reasoning effort and benchmark version are part of the comparison.
SpaceXAI release table, September 21, 2026. Vendor-reported comparison; model efforts differ as labeled. Important: Grok 4.7 uses high effort here, not xhigh.
SpaceXAI release table, September 21, 2026. Vendor-reported comparison; model efforts differ as labeled. Grok 4.7 has the highest score in this table.
SpaceXAI release table, September 21, 2026. Vendor-reported comparison; model efforts differ as labeled. Elo scores. These are the values in the vendor table, not an independently rerun evaluation.
SpaceXAI release table, September 21, 2026. Vendor-reported comparison; model efforts differ as labeled. Grok gains 17.7 points over 4.6. Fable 5.1 remains clearly ahead; do not mix these results with older Terminal-Bench versions.
SpaceXAI release table, September 21, 2026. Vendor-reported comparison; model efforts differ as labeled. Leading this displayed benchmark is not a substitute for professional review of legal work.
SpaceXAI release table, September 21, 2026. Vendor-reported comparison; model efforts differ as labeled. Grok improves over 4.6 but trails Fable and Sol. Benchmark performance does not establish suitability for autonomous clinical decisions.
SpaceXAI interactive launch chart, September 21, 2026. Elo scores. This chart compares GPT-6 Astra, unlike the release table above, which compares GPT-5.6 Sol.
SpaceXAI release-table prices. Lower is better. Grok 4.7 retains the same base rate as 4.6. Long-context, cached-input and priority pricing are separate.
SpaceXAI release-table prices. Lower is better. Token rate is not total task cost: reasoning tokens, context length and the number of attempts still matter.
Source: SpaceXAI's September 21, 2026 Grok 4.7 launch report and current xAI API documentation. All charts above are attributed vendor-reported results, not SPIRITT reruns. The release table uses GPT-5.6 Sol; GDPval uses GPT-6 Astra. Most Grok 4.7 scores use xhigh effort, while DeepSWE uses high. Standard prices shown exclude long-context and priority premiums. Full sources are linked below.
Give the work a proper home: files, tools, a browser and a concrete finish line.
Bring your repository, files, research or idea into SPIRITT. Your agent has a browser, terminal, integrations and memory, so the task can move beyond a conversation.

Use the release evidence to judge the fit: sustained coding, document work, verification loops and 500K context. Choose the model and access route available in your workspace before committing a real workflow.

Set a clear outcome and ask your agent to test its work. Keep the source files, results and decisions together, with human review wherever the stakes demand it.

Give your AI agent a workspace with the tools, files and memory to reach a real finish line.