GDPval-AA v2 Elo
Artificial Analysis same-harness result. Grok ranks third behind Opus 5 effort variants; its confidence interval overlaps Fable 5 and Qwen3.8 Max.
SpaceXAI tuned Grok 4.6 for long-running agents, difficult coding and knowledge work, and stronger first passes on interactive and visual projects. Independent Artificial Analysis puts it at 61 on the Intelligence Index, with frontier-level agent performance at $2/M input and $6/M output. On SPIRITT, it gets the browser, terminal, files, tools, and durable context to turn that persistence into finished work.
The official release, independent measurements, hands-on coding results, and an honest reminder that value does not mean best at everything.
Introducing Grok 4.6. It delivers frontier intelligence and is a significant improvement over Grok 4.5 at the same price.
— SpaceXAI (@SpaceXAI) August 12, 2026
Grok 4.6 is a significant jump in UI design. This one-shot landing page shows it in action, and the model is now live on SPIRITT in an agentic environment through the AI gateway or your own provider connection.
— Tamir (@TamirSPIRITT) August 13, 2026
SpaceXAI's Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index, joining the frontier in line with GPT-5.6 Sol, with standout agentic performance at lower cost.
— Artificial Analysis (@ArtificialAnlys) August 12, 2026
Grok 4.6 made large gains on AA-Briefcase, our agentic knowledge work benchmark, and cost substantially less than other leading models.
— Artificial Analysis (@ArtificialAnlys) August 12, 2026
Grok 4.6 is here! It landed in Code Arena: WebDev at #7 with 1618 points, a big jump from Grok 4.5 at #13, and on par with GPT-5.6 Sol xHigh and Claude Fable 5 as confidence intervals tighten.
— Arena.ai (@arena) August 12, 2026
Grok 4.6 is now available in Devin. Cognition reports a significant improvement over Grok 4.5, surpassing GPT-5.6 Sol and ranking behind only Opus 5 and Fable 5 in its evaluation.
— Cognition (@cognition) August 12, 2026
Grok 4.6 is live. A field guide shares learnings, findings, and tips for using the model.
— Eric Zakariasson (@ericzakariasson) August 12, 2026
Tried Grok 4.6 on my bug bench: 105 hidden bugs in two real repos, judged blind. Grok 4.5 found 17, Grok 4.6 found 27, and Fable 5 found 29. It may be my new default for time, value, and quality.
— Paweł Huryn (@PawelHuryn) August 12, 2026
Grok 4.6 is great. Its intelligence per dollar is crazy good. Fable 5 is still the best model, but SpaceXAI and Cursor are now firmly in the frontier intelligence game.
— Mckay Wrigley (@mckaywrigley) August 13, 2026
Grok 4.6 keeps the 500K context and headline price of Grok 4.5, then pushes harder on sustained execution: more agentic training, better self-checking on long trajectories, and more substantial visual and interactive first passes.
SpaceXAI trained Grok 4.6 across knowledge work, general coding, web development, CAD, kernel optimization, and other agentic environments. The launch report says longer runs increasingly showed the model testing and verifying its own work before moving on.
The API accepts text and images and returns text, supports low through xhigh reasoning effort, function calling and structured outputs, and has a 500,000-token context window. Standard pricing is $2/M input, $0.50/M cached input, and $6/M output; prompts at or above 200K move to $4/$1/$12 for the full request.
Artificial Analysis, August 12, 2026. Grok 4.6 high reaches the frontier at 61, tied with GPT-5.6 Sol max and behind Anthropic's two leaders. It is not the universal #1 model.
The strongest case is the combination: near-top intelligence, sustained agent work, practical speed, and much lower operating cost. The coding boards also show where Sol and Fable still lead.
Artificial Analysis same-harness result. Grok ranks third behind Opus 5 effort variants; its confidence interval overlaps Fable 5 and Qwen3.8 Max.
Artificial Analysis multi-turn customer-service tool-use evaluation. Grok is second in the displayed frontier pair.
Artificial Analysis: 88.4%, in line with the leaders on v2.1. Do not compare this directly with the much harder v3.0 row below.
Grok sits at Fable-tier but behind Opus 5. It averaged about 53 turns and 0.5B input tokens versus Opus 5 max at about 103 turns and 2.0B.
SpaceXAI launch table. Grok improves sharply over 4.5, but Sol and Fable still lead this long-horizon software-engineering benchmark.
SpaceXAI launch table, harder v3.0 task set. Grok improves 10.3 points over 4.5 but remains clearly behind Sol and Fable.
Lower is better. Grok delivers 61 Index points at $0.84 per evaluated task, putting it on Artificial Analysis's price-performance Pareto frontier.
Artificial Analysis model-level snapshots, August 13, 2026. Figures represent first-party API performance and may change with the rolling measurement window.
Sources: Artificial Analysis Grok 4.6 analysis and live model/provider pages; SpaceXAI's August 12 launch report, API documentation, and model card; Arena.ai, Cognition, and verified direct X posts. Independent and vendor-reported results are labeled separately. SpaceXAI's competitor rows use best self-reported or public results rather than one controlled rerun. Grok 4.6 is compelling for supervised, cost-sensitive agent work; use your own evals for consequential or low-oversight workflows.
From a fresh workspace to Grok 4.6 staying with the task across tools, files, and verification loops
Open a SPIRITT workspace with the browser, terminal, files, integrations, and memory already available. Grok 4.6's long-trajectory training matters when it can inspect evidence and act, not only answer in a chat box.

Choose Grok 4.6 in the model picker. Use high for the documented frontier operating point, then raise or lower reasoning effort when the consequence, latency, and cost of the task justify it.

Give it the objective, source material, constraints, and acceptance checks. Grok can research, code, use tools, and iterate through verification steps; you keep approvals and final judgment where mistakes would be expensive.

A full cloud workspace, browser, terminal, tools, files, and memory. Pick Grok 4.6 and give it a finish line.