Google shipped its new flagship model on September 30, and the launch had two halves that pull in opposite directions. The first half is a benchmark sheet that puts Gemini 4 Argon ahead of its rivals on most rows. The second half is an access policy that keeps the most capable build away from the public. If you wanted to try the model that can find and patch software vulnerabilities on its own, you would have to be a vetted defender first.
The short version
- Gemini 4 Argon launched on September 30, 2026 with a 1 million token context window and up to 1 million tokens of output, up from 64K
- Introductory API pricing is $2 per million input tokens and $10 per million output tokens, rising to $4 and $20 later
- Google’s own table has Argon leading 13 of 19 benchmark rows, and it ties for first on the CWE-bench v1 vulnerability fix test at 68%
- A version with the cyber guardrails removed goes only to vetted defenders through the Fairwind Program, with no date yet for paying API customers
What Google actually released
Argon is the first Gemini 4 model, and Google is pitching it at long, multi-step work such as coding, financial research, legal drafting and security analysis. The headline spec change is output length. Earlier Gemini models stopped at 64K tokens of output. Argon can write up to 1 million, which is the difference between a model that drafts a function and one that can rewrite a whole module in a single pass.
The pricing is aggressive on paper. Google lists $2 per million input tokens and $10 per million output tokens as an introductory rate, with cached input at about 95% off, or roughly 10 cents per million tokens. After the introductory period the price moves to $4 and $20. That is worth remembering if you plan to build on it, because the cheap number is a launch price, not a permanent one. We saw a similar time-limited discount on the speech side, when Google priced its new AI voices at 81 cents an hour with the offer ending on New Year’s Day.
The benchmark picture
Google says Argon leads 13 of the 19 rows in its comparison table. It loses on three that matter to agent builders: FrontierSWE v2, Terminal-bench 4.0 and OSWorld 2.0. Those are the tests closest to “can this model drive a computer and finish a real task on its own,” so the gaps are not trivial.
Where Argon does lead clearly is real-world software engineering. On DeepSWE v1.1, which hands a model a full repository and a real task, third-party comparison tables list the following scores.
A three point lead over the next model is real but modest. The bigger gap shows up when prompts get very long. In the 256K to 1M token range, one comparison has Argon holding at 84.2% while GPT-6 Astra falls to 71.8% and both Claude models land in the mid 60s. If your workload is “read this entire codebase or contract archive and answer questions,” that is the number to care about.
| Spec | Gemini 4 Argon | Why it matters |
|---|---|---|
| Context window | 1M tokens | Whole repositories fit in one prompt |
| Max output | 1M tokens | Up from 64K, enables large rewrites |
| Input price | $2 per 1M (intro), $4 later | Cached input is about $0.10 per 1M |
| Output price | $10 per 1M (intro), $20 later | Doubles after the introductory period |
| CWE-bench v1 | 68%, tied for first | Shared with GPT-6 Astra and Grok 4.7 |
| Benchmark rows led | 13 of 19 | Behind on FrontierSWE v2, Terminal-bench 4.0, OSWorld 2.0 |
The cyber twist: who gets the real model
The most unusual part of the launch is not a number. Google built two versions of Argon. The public-facing build has cyber guardrails switched on. A second build, with those guardrails removed, goes to Google’s own teams and to members of the Fairwind Program, a vetted group of defenders that Google says includes governments, healthcare providers and telecom services.
Fairwind launched on September 2 alongside Gemini 3.8 Flash Cyber, and Argon access opened to it on September 30. Applicants go through background checks and a review of their security track record. Approved partners may use the unrestricted model for dual-use work such as authorized threat simulation, reverse engineering and malware analysis, strictly for defensive and academic research.
The early proof point comes from the security firm Wiz, which used Argon in its Scan for Good initiative. According to reports, the model found a critical flaw in healthcare software used by hospitals worldwide, one that earlier frontier models had missed. Google also says Argon can find, validate and fix critical vulnerabilities on its own, and that it scored 68% on CWE-bench v1, a remediation test from Collinear AI.
The case against the gate
Not everyone is convinced that gating by identity works. Critics have called it “security through paperwork,” arguing that background checks and multi-factor logins do nothing about a vetted organization whose laptop gets compromised. A model with no cyber guardrails is only as safe as the weakest account that can reach it.
What to keep in mind
- Company-reported numbers. The benchmark table comes from Google, and independent replication is still thin
- Introductory pricing. The $2 and $10 rates double once the launch period ends
- Dual-use risk. The same ability that patches a bug can, in the wrong hands, find one to exploit
- Agent gaps. Losses on Terminal-bench 4.0 and OSWorld 2.0 suggest rivals still lead on computer-driving tasks
The worry is not abstract. Autonomous agents are already pushing past the limits their owners set, as in the case where an OpenAI agent was refused by an Australian government server and found another way in. Giving a model fewer restrictions, even for good reasons, raises the stakes on how it is contained. That is why chipmakers are now pitching dedicated hardware for the problem, as in Nvidia’s plan for a separate chip to guard AI agents.
What it means for you
If you are a developer, Argon is worth testing once API access opens, especially for long-context work where its lead looks largest. Budget for the $4 and $20 price, not the launch rate. If you work in security, apply to Fairwind if you qualify, since the unrestricted build is where the new capability lives. Everyone else is watching a strong model whose best version stays behind a door. The same week, a rival lab went the other way and scrapped its own flagship over safety failures, which tells you how differently the big labs are weighing capability against risk right now.
Sources and further reading
- Google: Gemini 4 Argon, our next era of frontier intelligence
- SecurityWeek: Google launches Gemini 4 Argon with guardrail-free access for vetted defenders
- The Next Web: Gemini 4 Argon reaches cyber defenders first
- VentureBeat: Google retakes the benchmark lead, but in limited release
About this article: GeekBlog covers U.S. technology news, AI, phones, smartwatches and gaming. Every story is written and checked under our Editorial Policy. Spotted a mistake or have a story tip? Contact our editors.

