Argon leads the Vals Index for professional work and pushes long-horizon coding forward. The broader benchmark picture shows both strengths and trade-offs.
Google’s next frontier model, in a phased rollout
Google announced Gemini 4 Argon on September 30, 2026, targeting software engineering, legal and financial work, multimodal research, and cybersecurity defense. Access starts with trusted testers and cyber defenders through the Fairwind Program. Wider access is planned, not generally available at announcement.
Google is expanding the output limit from the previous 64K to up to one million tokens. Artificial Analysis tested Argon at high reasoning with Long Decode Continuation enabled. A larger output budget creates room for extended reasoning and generation; it is not a speed or low-cost guarantee.
Argon scores 68.9% on Vals Index v2.1, leading its professional-work evaluation. On Artificial Analysis’s broader Intelligence Index v4.3.2, it scores 53, matching the rounded score of GPT-6 Astra at max effort while trailing the leading Claude configurations.
Introductory API pricing will be $2 per million input tokens and $10 per million output tokens. Google says $4/$20 rates apply after the introductory period, without announcing its end date. Cached input is discounted by 95%. These are Google API prices, not SPIRITT subscription prices.
Inside Google, Argon has helped with large code migrations, memory optimization and specialized research. Google reports memory optimizations that would free over 300 TiB once rolled out, and a 2.7× speedup versus an existing Rust video-decoder port. Critical code rewrites are still subject to extensive auditing and testing; these are specific internal examples, not universal productivity promises.
Up to 1M output tokensHigh-effort reasoningCoding + knowledge work$2 / $10 introductory API ratesPhased rolloutRelease brief
Artificial Analysis Intelligence Index v4.3.2
Claude Opus 5.5 (max with fallback)
58
Claude Sonnet 5.5 (max with fallback)
56
Claude Fable 5.1 (max with fallback)
53
Artificial Analysis, September 30, 2026, Index v4.3.2. Selected configurations; rounded index points, not percentages. Argon uses high reasoning with Long Decode Continuation, up to 1M output tokens. Claude fallback and effort settings are shown. Matching rounded scores do not establish equal capability on every task.
Independent professional work
WinnerVals Index v2.1
Vals AI’s September 30 launch snapshot. GDP-weighted finance, coding, legal and tax tasks. Argon: high reasoning, temperature 1, 262K output budget. Comparator effort settings are not specified in the captured chart. This is a different evaluation from the AA Intelligence Index.
Business workflow execution
WinnerAutomationBench: Zapier
Google’s comparison, citing Zapier’s private-set leaderboard. These scores are separate from AutomationBench-AA’s objective-completion metric. Model settings and safeguards follow the cited runs, not a uniform rerun.
Multi-step financial research
WinnerVals Finance Agent v2
Google’s comparison, citing Vals AI. Tests financial research that requires multiple steps and reading source documents. The displayed model settings follow Google’s evaluation report.
Legal research and drafting
WinnerHarvey’s Legal Agent Benchmark
Google’s comparison, citing Vals AI. Argon leads the listed systems, but a 19.6% score still leaves substantial room to improve. Relative leadership is not reliable completion of every legal task.
Long-horizon software work
WinnerDeepSWE v1.1
Google computes Argon with a mini-swe-agent harness; Astra comes from the public leaderboard, and Claude values from system cards. Highest-scoring reported thinking levels are selected. These are compiled results, not one shared rerun.
Difficult software engineering
FrontierSWE v2
Google’s comparison, citing Proximal’s leaderboard. Argon trails all three displayed peers here. Strength on DeepSWE does not establish a lead on every coding evaluation.
Vibe Code Bench
Google’s comparison, citing Vals AI’s leaderboard. All four listed systems score highly; the small point-estimate gaps do not establish a decisive advantage for every app-building task.
Terminal agents: Google compilation
Terminal-Bench 4.0: Google
Argon is Google’s run; other scores come from the benchmark leaderboard at each model’s highest-scoring reported thinking setting. Argon trails the listed peers. Keep this separate from AA’s mini-swe-agent results.
Machine-learning engineering
PostTrainBench v1.1
Google runs all models with OpenCode, ten hours and one NVIDIA H100 per attempt. Weighted aggregate across four base models and seven benchmarks. Argon trails Opus 5.5 in this setup.
Scientific work in a terminal
Terminal-Bench Science 0.1
Google’s Argon run uses a 6× verifier timeout to address verification timeouts; other scores come from the public leaderboard. Argon trails Astra and Opus. The timeout difference matters when comparing results.
Biological research tasks
WinnerLABBench 2
Google evaluates all four models with a Linux terminal, bioinformatics tools, Python, R and internet access. These are tool-assisted results, not closed-book biological knowledge scores.
Advanced mathematics
WinnerRiemannBench
Google’s comparison, citing Surge’s public leaderboard. Argon has the highest displayed score in this snapshot; general reasoning ability is broader than one mathematics evaluation.
Reasoning over shorter graph contexts
GraphWalks: up to 128K
Google-run breadth-first-search F1 on 650 items with contexts up to 128K tokens. The 99.7% and 98.7% scores are close. F1 is not a whole-task completion rate.
Reasoning over long graph contexts
WinnerGraphWalks: 256K to 1M
Google-run breadth-first-search F1 on 200 problems with contexts from 256K to 1M tokens. This evaluates graph reasoning at those lengths; it is not a complete specification of the future public API context window.
Long-running computer use
Agent’s Last Exam
Google’s Argon run: default ALE-Claw, five-hour window, safeguards enabled, binary task pass rate. Astra and Opus come from the public leaderboard. Fable has no reported comparable value and is omitted.
Computer-use progress, not full completion
OSWorld 2.0: offline partial credit
Google reports the best of three Argon runs, not an ordinary average: 1080p screenshots, 500-step limit, compaction, parallel tools and safeguards. Astra is from its launch report. Partial credit on the offline subset is not full-task success; incompatible Claude combined-subset scores are omitted.
Understanding charts without tools
Chartography
Google’s comparison, citing Surge’s leaderboard. No tools. Argon and Astra are close on this chart-understanding evaluation; the gap alone is not evidence of a decisive lead.
Understanding long videos
WinnerLVBench
Google evaluates all four models without tools. Gemini uses 1 frame/second; Astra, Fable and Opus are capped at 800, 300 and 600 frames respectively because of API limits. Unequal sampling limits direct comparability.
Security vulnerability remediation
Near frontierCWE-bench v1
Gemini 4 Argon (Antigravity)
68%
Claude Opus 5.5 (Claude Code)
67%
Claude Fable 5.1 (Claude Code)
58%
Selected leading models from the leaderboard reproduced by Google. Argon, Grok 4.7 and Astra tie at 68% pass@1; pass@4 breaks leaderboard ties. Different agent harnesses are shown, so this is not a unique first-place claim.
Google’s internal defensive evaluation
WinnerReal-world vulnerability discovery
Gemini 3.8 Flash Cyber
71%
Google’s private dataset measures recall of recent confirmed vulnerabilities across popular open-source projects in 20 languages. Models receive source code through an internal Antigravity harness. These published scores are not an independently reproduced public benchmark.
Wiz’s internal defensive evaluation
WinnerWiz penetration-testing benchmark
Gemini 3.8 Flash Cyber
58.2%
Google reports Wiz’s private benchmark of real web vulnerabilities without source-code access. Published internal-evaluation scores, not a general vulnerability-detection guarantee or an independently reproduced public benchmark.
Resistance to indirect prompt injection
WinnerGray Swan IPI: 15 attack attempts
Google reports Gray Swan’s attack-success rates at 15 attempts; lower is better. Selected peers shown. Per-model effort and harness details are not specified here. Low measured success does not mean immunity to prompt injection.
Independent business automation
AutomationBench-AA: objectives
Gemini 4 Argon (high)
77.5%
Artificial Analysis: dataset 1.0.6, 657 private tasks, 50-turn cap. Average share of guardrail-adjusted objectives completed, with errors or guardrail violations scoring zero. Different metric and setup from Zapier’s 51.3% result; not a whole-task pass rate.
Independent terminal-agent run
Terminal-Bench 4.0: AA
Artificial Analysis’s rounded launch scores: mini-swe-agent, 66 tasks, three repeats, 500-step cap, no compaction. Claude configurations use the evaluator’s default fallbacks. These scores are separate from Google’s compiled results and Vals’ 57.6% run.
Introductory cost advantage
AA cost per Index task
Gemini 4 Argon: introductory
$1.99
Gemini 4 Argon: standard
$3.98
Artificial Analysis’s task-cost estimate, not cost per successful completion. Argon high uses about 62K output tokens per task versus Astra max at 27K. The introductory advantage comes from pricing; after the promotion, Argon costs more in this comparison. These are not Vals task costs.
Fewer unsupported answers
AA-Omniscience: hallucination metric
Artificial Analysis’s benchmark-specific hallucination metric, lower is better. Argon also answers fewer questions correctly: see the adjacent accuracy comparison. More abstention can reduce wrong guesses without adding knowledge; this is not a real-world error rate.
AA-Omniscience: accuracy
Artificial Analysis’s rounded share of questions answered correctly. Argon’s lower hallucination metric does not translate to higher accuracy here. Its overall Omniscience score is 42 versus Astra’s 43.
Announced API token prices
Argon input: per million tokens
After introductory period
$4.00
Google’s announced USD API rates. The promotion end date is not specified, and general access is still rolling out. Cached input is 95% off the applicable input rate. Ancillary charges and product-specific terms are not fully specified; these are not SPIRITT plan prices.
Announced API token prices
Argon output: per million tokens
After introductory period
$20.00
Google’s announced USD API rates. The promotion end date is not specified, and general access is still rolling out. Cached input is 95% off the applicable input rate. Ancillary charges and product-specific terms are not fully specified; these are not SPIRITT plan prices.
Sources: Google’s September 30, 2026 announcement and evaluation methodology; Artificial Analysis Intelligence Index v4.3.2 and launch results; Vals Index v2.1. Google’s comparison table combines its own runs with evaluator and provider-reported results. Versions, reasoning settings, tools, safeguards and scoring differ. Missing scores are omitted, not plotted as zero. Close point estimates do not establish a meaningful advantage. API token prices, evaluation task costs and SPIRITT subscription prices are different things.