SPIRITT logoSPIRITTGoogle DeepMindGemini 4 Argon

Gemini 4 Argon. Built for the hardest work.

Complex codebases. Deep research. Multi-step business work. Google’s new frontier model brings up to 1M output tokens to longer problems. Explore the early results, introductory pricing and phased rollout.

The Argon launch, through different lenses.

Google’s announcement, independent evaluations, visual benchmark examples and early reactions. Access is still limited; these are launch and evaluation posts, not a gallery of public hands-on demos.

More room to reason. Bigger problems to solve.

Argon leads the Vals Index for professional work and pushes long-horizon coding forward. The broader benchmark picture shows both strengths and trade-offs.

Google’s next frontier model, in a phased rollout

Google announced Gemini 4 Argon on September 30, 2026, targeting software engineering, legal and financial work, multimodal research, and cybersecurity defense. Access starts with trusted testers and cyber defenders through the Fairwind Program. Wider access is planned, not generally available at announcement.

Google is expanding the output limit from the previous 64K to up to one million tokens. Artificial Analysis tested Argon at high reasoning with Long Decode Continuation enabled. A larger output budget creates room for extended reasoning and generation; it is not a speed or low-cost guarantee.

Argon scores 68.9% on Vals Index v2.1, leading its professional-work evaluation. On Artificial Analysis’s broader Intelligence Index v4.3.2, it scores 53, matching the rounded score of GPT-6 Astra at max effort while trailing the leading Claude configurations.

Introductory API pricing will be $2 per million input tokens and $10 per million output tokens. Google says $4/$20 rates apply after the introductory period, without announcing its end date. Cached input is discounted by 95%. These are Google API prices, not SPIRITT subscription prices.

Inside Google, Argon has helped with large code migrations, memory optimization and specialized research. Google reports memory optimizations that would free over 300 TiB once rolled out, and a 2.7× speedup versus an existing Rust video-decoder port. Critical code rewrites are still subject to extensive auditing and testing; these are specific internal examples, not universal productivity promises.

Up to 1M output tokensHigh-effort reasoningCoding + knowledge work$2 / $10 introductory API ratesPhased rolloutRelease brief

Artificial Analysis Intelligence Index v4.3.2

Claude Opus 5.5 (max with fallback)
58
Claude Sonnet 5.5 (max with fallback)
56
Claude Fable 5.1 (max with fallback)
53
GPT-6 Astra (max)
53
Gemini 4 Argon (high)
53
GPT-6.1 Sol (max)
52
Grok 4.7 (xhigh)
46

Artificial Analysis, September 30, 2026, Index v4.3.2. Selected configurations; rounded index points, not percentages. Argon uses high reasoning with Long Decode Continuation, up to 1M output tokens. Claude fallback and effort settings are shown. Matching rounded scores do not establish equal capability on every task.

Where Argon leads. Where the trade-offs are.

Professional work, coding, science, long context and multimodal understanding, with independent evaluations kept separate from Google’s compiled results. Compare the setup as well as the score.

Independent professional work
Winner

Vals Index v2.1

Gemini 4 Argon
68.9%
Claude Sonnet 5.5
67.04%
Claude Opus 5.5
66.97%
Claude Fable 5.1
65.83%
GPT-6 Astra
63.13%
GPT-6.1 Sol
61.15%

Vals AI’s September 30 launch snapshot. GDP-weighted finance, coding, legal and tax tasks. Argon: high reasoning, temperature 1, 262K output budget. Comparator effort settings are not specified in the captured chart. This is a different evaluation from the AA Intelligence Index.

Business workflow execution
Winner

AutomationBench: Zapier

Gemini 4 Argon
51.3%
Claude Opus 5.5
42.5%
GPT-6 Astra
41.4%
Claude Fable 5.1
31.4%

Google’s comparison, citing Zapier’s private-set leaderboard. These scores are separate from AutomationBench-AA’s objective-completion metric. Model settings and safeguards follow the cited runs, not a uniform rerun.

Multi-step financial research
Winner

Vals Finance Agent v2

Gemini 4 Argon
65.4%
Claude Fable 5.1
58.9%
Claude Opus 5.5
58.6%
GPT-6 Astra
53.5%

Google’s comparison, citing Vals AI. Tests financial research that requires multiple steps and reading source documents. The displayed model settings follow Google’s evaluation report.

Legal research and drafting
Winner

Harvey’s Legal Agent Benchmark

Gemini 4 Argon
19.6%
Claude Fable 5.1
6.7%
GPT-6 Astra
5.4%
Claude Opus 5.5
3.8%

Google’s comparison, citing Vals AI. Argon leads the listed systems, but a 19.6% score still leaves substantial room to improve. Relative leadership is not reliable completion of every legal task.

Long-horizon software work
Winner

DeepSWE v1.1

Gemini 4 Argon
77.9%
Claude Opus 5.5
74.2%
GPT-6 Astra
74.1%
Claude Fable 5.1
67.4%

Google computes Argon with a mini-swe-agent harness; Astra comes from the public leaderboard, and Claude values from system cards. Highest-scoring reported thinking levels are selected. These are compiled results, not one shared rerun.

Difficult software engineering

FrontierSWE v2

GPT-6 Astra
65.5%
Claude Opus 5.5
62.3%
Claude Fable 5.1
56.3%
Gemini 4 Argon
55%

Google’s comparison, citing Proximal’s leaderboard. Argon trails all three displayed peers here. Strength on DeepSWE does not establish a lead on every coding evaluation.

Building from a brief

Vibe Code Bench

Gemini 4 Argon
91.9%
Claude Fable 5.1
90.3%
Claude Opus 5.5
90.3%
GPT-6 Astra
89.6%

Google’s comparison, citing Vals AI’s leaderboard. All four listed systems score highly; the small point-estimate gaps do not establish a decisive advantage for every app-building task.

Terminal agents: Google compilation

Terminal-Bench 4.0: Google

Claude Opus 5.5
66.4%
GPT-6 Astra
58.2%
Claude Fable 5.1
57.9%
Gemini 4 Argon
57.4%

Argon is Google’s run; other scores come from the benchmark leaderboard at each model’s highest-scoring reported thinking setting. Argon trails the listed peers. Keep this separate from AA’s mini-swe-agent results.

Machine-learning engineering

PostTrainBench v1.1

Claude Opus 5.5
49.3%
Gemini 4 Argon
45.3%
GPT-6 Astra
44.3%
Claude Fable 5.1
40.2%

Google runs all models with OpenCode, ten hours and one NVIDIA H100 per attempt. Weighted aggregate across four base models and seven benchmarks. Argon trails Opus 5.5 in this setup.

Scientific work in a terminal

Terminal-Bench Science 0.1

GPT-6 Astra
68.1%
Claude Opus 5.5
63.3%
Gemini 4 Argon
57.6%
Claude Fable 5.1
52.6%

Google’s Argon run uses a 6× verifier timeout to address verification timeouts; other scores come from the public leaderboard. Argon trails Astra and Opus. The timeout difference matters when comparing results.

Biological research tasks
Winner

LABBench 2

Gemini 4 Argon
88.8%
GPT-6 Astra
85.4%
Claude Opus 5.5
73.1%
Claude Fable 5.1
68.6%

Google evaluates all four models with a Linux terminal, bioinformatics tools, Python, R and internet access. These are tool-assisted results, not closed-book biological knowledge scores.

Advanced mathematics
Winner

RiemannBench

Gemini 4 Argon
76%
GPT-6 Astra
72%
Claude Opus 5.5
69.6%
Claude Fable 5.1
65.6%

Google’s comparison, citing Surge’s public leaderboard. Argon has the highest displayed score in this snapshot; general reasoning ability is broader than one mathematics evaluation.

Reasoning over shorter graph contexts

GraphWalks: up to 128K

Gemini 4 Argon
99.7%
GPT-6 Astra
98.7%
Claude Fable 5.1
91.4%
Claude Opus 5.5
90.6%

Google-run breadth-first-search F1 on 650 items with contexts up to 128K tokens. The 99.7% and 98.7% scores are close. F1 is not a whole-task completion rate.

Reasoning over long graph contexts
Winner

GraphWalks: 256K to 1M

Gemini 4 Argon
84.2%
GPT-6 Astra
71.8%
Claude Opus 5.5
66.8%
Claude Fable 5.1
65%

Google-run breadth-first-search F1 on 200 problems with contexts from 256K to 1M tokens. This evaluates graph reasoning at those lengths; it is not a complete specification of the future public API context window.

Long-running computer use

Agent’s Last Exam

Gemini 4 Argon
39.5%
Claude Opus 5.5
38.2%
GPT-6 Astra
34.2%

Google’s Argon run: default ALE-Claw, five-hour window, safeguards enabled, binary task pass rate. Astra and Opus come from the public leaderboard. Fable has no reported comparable value and is omitted.

Computer-use progress, not full completion

OSWorld 2.0: offline partial credit

GPT-6 Astra
72.6%
Gemini 4 Argon
69.2%

Google reports the best of three Argon runs, not an ordinary average: 1080p screenshots, 500-step limit, compaction, parallel tools and safeguards. Astra is from its launch report. Partial credit on the offline subset is not full-task success; incompatible Claude combined-subset scores are omitted.

Understanding charts without tools

Chartography

Gemini 4 Argon
71.6%
GPT-6 Astra
71%
Claude Opus 5.5
66.3%
Claude Fable 5.1
46.2%

Google’s comparison, citing Surge’s leaderboard. No tools. Argon and Astra are close on this chart-understanding evaluation; the gap alone is not evidence of a decisive lead.

Understanding long videos
Winner

LVBench

Gemini 4 Argon
91.7%
GPT-6 Astra
87.5%
Claude Opus 5.5
83.7%
Claude Fable 5.1
79.7%

Google evaluates all four models without tools. Gemini uses 1 frame/second; Astra, Fable and Opus are capped at 800, 300 and 600 frames respectively because of API limits. Unequal sampling limits direct comparability.

Security vulnerability remediation
Near frontier

CWE-bench v1

Grok 4.7 (opencode)
68%
Gemini 4 Argon (Antigravity)
68%
GPT-6 Astra (Codex)
68%
Claude Opus 5.5 (Claude Code)
67%
Claude Fable 5.1 (Claude Code)
58%

Selected leading models from the leaderboard reproduced by Google. Argon, Grok 4.7 and Astra tie at 68% pass@1; pass@4 breaks leaderboard ties. Different agent harnesses are shown, so this is not a unique first-place claim.

Google’s internal defensive evaluation
Winner

Real-world vulnerability discovery

Gemini 4 Argon
85.8%
Gemini 3.8 Flash Cyber
71%

Google’s private dataset measures recall of recent confirmed vulnerabilities across popular open-source projects in 20 languages. Models receive source code through an internal Antigravity harness. These published scores are not an independently reproduced public benchmark.

Wiz’s internal defensive evaluation
Winner

Wiz penetration-testing benchmark

Gemini 4 Argon
70.9%
Gemini 3.8 Flash Cyber
58.2%

Google reports Wiz’s private benchmark of real web vulnerabilities without source-code access. Published internal-evaluation scores, not a general vulnerability-detection guarantee or an independently reproduced public benchmark.

Resistance to indirect prompt injection
Winner

Gray Swan IPI: 15 attack attempts

Gemini 4 Argon
0.7%
Claude Opus 5.5
1%
Claude Fable 5.1
1%
Gemini 3.8 Flash Cyber
6%
GPT-6 Astra
8.5%
Grok 4.6
51.8%

Google reports Gray Swan’s attack-success rates at 15 attempts; lower is better. Selected peers shown. Per-model effort and harness details are not specified here. Low measured success does not mean immunity to prompt injection.

Independent business automation

AutomationBench-AA: objectives

Gemini 4 Argon (high)
77.5%
Sonnet 5.5 (max)
71.3%

Artificial Analysis: dataset 1.0.6, 657 private tasks, 50-turn cap. Average share of guardrail-adjusted objectives completed, with errors or guardrail violations scoring zero. Different metric and setup from Zapier’s 51.3% result; not a whole-task pass rate.

Independent terminal-agent run

Terminal-Bench 4.0: AA

Sonnet 5.5 (max)
64%
Opus 5.5 (max)
60%
GPT-6 Astra (max)
59%
Gemini 4 Argon (high)
57%

Artificial Analysis’s rounded launch scores: mini-swe-agent, 66 tasks, three repeats, 500-step cap, no compaction. Claude configurations use the evaluator’s default fallbacks. These scores are separate from Google’s compiled results and Vals’ 57.6% run.

Introductory cost advantage

AA cost per Index task

Gemini 4 Argon: introductory
$1.99
GPT-6 Astra (max)
$3.26
Gemini 4 Argon: standard
$3.98

Artificial Analysis’s task-cost estimate, not cost per successful completion. Argon high uses about 62K output tokens per task versus Astra max at 27K. The introductory advantage comes from pricing; after the promotion, Argon costs more in this comparison. These are not Vals task costs.

Fewer unsupported answers

AA-Omniscience: hallucination metric

Gemini 4 Argon (high)
15%
GPT-6 Astra (max)
51%
GPT-6.1 Sol (max)
54%

Artificial Analysis’s benchmark-specific hallucination metric, lower is better. Argon also answers fewer questions correctly: see the adjacent accuracy comparison. More abstention can reduce wrong guesses without adding knowledge; this is not a real-world error rate.

The accuracy trade-off

AA-Omniscience: accuracy

GPT-6 Astra (max)
63%
Gemini 4 Argon (high)
50%

Artificial Analysis’s rounded share of questions answered correctly. Argon’s lower hallucination metric does not translate to higher accuracy here. Its overall Omniscience score is 42 versus Astra’s 43.

Announced API token prices

Argon input: per million tokens

Introductory price
$2.00
After introductory period
$4.00

Google’s announced USD API rates. The promotion end date is not specified, and general access is still rolling out. Cached input is 95% off the applicable input rate. Ancillary charges and product-specific terms are not fully specified; these are not SPIRITT plan prices.

Announced API token prices

Argon output: per million tokens

Introductory price
$10.00
After introductory period
$20.00

Google’s announced USD API rates. The promotion end date is not specified, and general access is still rolling out. Cached input is 95% off the applicable input rate. Ancillary charges and product-specific terms are not fully specified; these are not SPIRITT plan prices.

Sources: Google’s September 30, 2026 announcement and evaluation methodology; Artificial Analysis Intelligence Index v4.3.2 and launch results; Vals Index v2.1. Google’s comparison table combines its own runs with evaluator and provider-reported results. Versions, reasoning settings, tools, safeguards and scoring differ. Missing scores are omitted, not plotted as zero. Close point estimates do not establish a meaningful advantage. API token prices, evaluation task costs and SPIRITT subscription prices are different things.

Give ambitious work a place to happen.

Bring the project into SPIRITT, start with your available models, and assess new releases against work that matters to you.

01

Bring the real problem

A codebase to improve, research to turn into a decision, or a workflow to automate. Give SPIRITT the files, context and success criteria so the work starts with a clear destination.

A soft glass illustration of a connected SPIRITT workspace
02

Follow Argon’s rollout

Google is starting with trusted testers before expanding access. Check the official rollout and the models or connections enabled in your workspace. This release brief does not mean Argon is already available in SPIRITT.

A multicolored glass Gemini star above translucent steps
03

Judge the finished result

Start with the models available to you today. When evaluating a new model, compare the quality of the finished work, total cost, time and human review needed, not just a leaderboard score.

A soft glass illustration of connected tools and a working automation

Bring the hard problem. Start the work.

Give SPIRITT your goal, files and context. Build with the models available in your workspace today, and evaluate new releases against real work as access becomes available.

Questions

Know what you’re choosing.

01What is Gemini 4 Argon?+
Gemini 4 Argon is Google’s frontier model announced on September 30, 2026. It focuses on sustained reasoning for complex software engineering, professional knowledge work, multimodal understanding and defensive cybersecurity. Google positions it as the start of a new Gemini generation.
02When can I use Gemini 4 Argon?+
At announcement, access starts with trusted testers and cyber defenders through Google’s Fairwind Program. Google plans a broader release beginning with paid API customers and Google AI Ultra subscribers after further safety work and feedback. It has not announced a firm date for general availability.
03What will Gemini 4 Argon cost?+
Google announced introductory API prices of $2 per million input tokens and $10 per million output tokens, followed by $4 and $20 after the introductory period. Cached input is 95% cheaper than the applicable input rate: $0.10 per million at the introductory rate. Google has not specified when the introductory period ends. These are API prices, not SPIRITT plan prices or a guaranteed cost for completing a task.
04Does one million output tokens mean a one-million-token context window?+
No. Output capacity and input context are different limits. Google’s headline announcement raises output capacity from 64K to up to 1M tokens for extended reasoning and generation. Artificial Analysis used Long Decode Continuation to test that capacity; Vals used a smaller 262K output budget. Do not treat the headline as a default response size or infer every API input/output setting from it.
05Is Argon the best model on every benchmark?+
No. It leads the Vals Index v2.1 at 68.9% for GDP-weighted professional work. Its Artificial Analysis Intelligence Index score is 53, behind the leading Claude configurations in the same snapshot. Google’s own comparison table shows Argon behind other models on FrontierSWE v2, Terminal-Bench 4.0, PostTrainBench and some computer-use or science results.
06Why are there different AutomationBench results?+
Google cites 51.3% from Zapier’s AutomationBench leaderboard. Artificial Analysis reports 77.5% on AutomationBench-AA, which measures the average share of guardrail-adjusted objectives completed. Different evaluation setups and scoring mean these are not interchangeable results or two measurements of the same completion rate.
07Why do the Terminal-Bench scores differ?+
Google reports 57.4% for Argon on Terminal-Bench 4.0, alongside scores compiled from other sources. Artificial Analysis’s separate mini-swe-agent evaluation reports a rounded 57%. Vals reports 57.6% in its own launch results. Harnesses, settings and rounding differ; one evaluator’s figure is not extra decimal precision for another evaluator’s run.
08Does Argon hallucinate less?+
On AA-Omniscience, Artificial Analysis reports a 15% hallucination metric for Argon versus 51% for GPT-6 Astra at max effort. But Argon’s accuracy is also lower: 50% versus 63%. Abstaining more can reduce wrong answers without answering more questions correctly. These are benchmark-specific measurements, not a real-world error rate or a guarantee of factual reliability.
09Can Argon understand images and video?+
Google reports results for chart understanding and long-video analysis, including 91.7% on LVBench. Those results establish evaluated multimodal capabilities, not every supported file format or launch API option. Video understanding is different from generating video; product-specific limits should be checked when wider access opens.
10What do Google’s internal coding examples prove?+
They show promising uses in specific engineering workflows, not guaranteed results for every project. Google reports memory optimizations that would free over 300 TiB once rolled out, ongoing C/C++ to Rust migrations, and a libgav1 rewrite that ran 2.7 times faster than an existing Rust port. The speedup was not measured against the optimized C++ implementation, and critical rewrites still require auditing, testing and human review.
11How is the cybersecurity rollout restricted?+
Google is providing selected trusted defenders and internal teams access without cyber guardrails through a controlled rollout. That is not a public unrestricted mode. Google says broader availability depends on strengthened safeguards against misuse, prompt injection and misalignment, alongside secure testing environments.
12How should I compare these benchmarks?+
Compare results from the same evaluator, version, scoring method and reasoning setup. A partial-credit computer-use score is not a full-task completion rate. A 1M output test is not equivalent to a 262K budget. Introductory task costs can change when pricing changes, and close scores do not establish a statistically meaningful lead.
13Can I use Gemini 4 Argon in SPIRITT?+
This page is a release brief, not an availability announcement. Argon access has not been confirmed for SPIRITT. Open a workspace to start with its available models and check the models and connections enabled for your account as Google’s rollout progresses.
Buy from builders who use what they sellBuilt usingSPIRITT