Close Menu
GeekBlog

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    What's Hot

    A Broken Test Tube Killed a Russian Lab Worker. The Official Cause of Death Keeps Changing.

    October 6, 2026

    Garmin Enduro 4 Is Official: Same 320-Hour Battery, Twice the Storage, and a Familiar $899.99 Price

    October 6, 2026

    North Korea Says Its New Missile Uses AI. The Proof It Offered Is a 20-Kilometer Disagreement.

    October 6, 2026
    Facebook
    GeekBlog
    • Home
    • Mobile
    • Tech News
    • Blog
    • Gaming
    • Smartwatch
    • How-To Guides
    • AI & Software
    Facebook
    GeekBlog
    Home»AI & Software»Google’s Gemini 4 Argon Tops Most Benchmarks. Almost Nobody Can Use It Yet.
    AI & Software

    Google’s Gemini 4 Argon Tops Most Benchmarks. Almost Nobody Can Use It Yet.

    Olivia HartmanBy Olivia HartmanOctober 5, 20267 Mins Read
    Share Facebook Twitter Pinterest LinkedIn Tumblr Email Copy Link
    Padlock resting on a computer keyboard, representing cybersecurity and Gemini 4 Argon cyber defense
    Photo: Pexels
    Share
    Facebook Twitter LinkedIn Pinterest Email Copy Link

    Google shipped its new flagship model on September 30, and the launch had two halves that pull in opposite directions. The first half is a benchmark sheet that puts Gemini 4 Argon ahead of its rivals on most rows. The second half is an access policy that keeps the most capable build away from the public. If you wanted to try the model that can find and patch software vulnerabilities on its own, you would have to be a vetted defender first.

    The short version

    • Gemini 4 Argon launched on September 30, 2026 with a 1 million token context window and up to 1 million tokens of output, up from 64K
    • Introductory API pricing is $2 per million input tokens and $10 per million output tokens, rising to $4 and $20 later
    • Google’s own table has Argon leading 13 of 19 benchmark rows, and it ties for first on the CWE-bench v1 vulnerability fix test at 68%
    • A version with the cyber guardrails removed goes only to vetted defenders through the Fairwind Program, with no date yet for paying API customers

    What Google actually released

    Argon is the first Gemini 4 model, and Google is pitching it at long, multi-step work such as coding, financial research, legal drafting and security analysis. The headline spec change is output length. Earlier Gemini models stopped at 64K tokens of output. Argon can write up to 1 million, which is the difference between a model that drafts a function and one that can rewrite a whole module in a single pass.

    The pricing is aggressive on paper. Google lists $2 per million input tokens and $10 per million output tokens as an introductory rate, with cached input at about 95% off, or roughly 10 cents per million tokens. After the introductory period the price moves to $4 and $20. That is worth remembering if you plan to build on it, because the cheap number is a launch price, not a permanent one. We saw a similar time-limited discount on the speech side, when Google priced its new AI voices at 81 cents an hour with the offer ending on New Year’s Day.

    The benchmark picture

    Google says Argon leads 13 of the 19 rows in its comparison table. It loses on three that matter to agent builders: FrontierSWE v2, Terminal-bench 4.0 and OSWorld 2.0. Those are the tests closest to “can this model drive a computer and finish a real task on its own,” so the gaps are not trivial.

    Recommended for you:

    The FTC Is Investigating OpenAI and Anthropic Over Rogue AI Agents. No Subpoenas Have Gone Out Yet.
    AI & Software·Oct 1, 2026

    The FTC Is Investigating OpenAI and Anthropic Over Rogue AI Agents. No Subpoenas Have Gone Out Yet.

    Where Argon does lead clearly is real-world software engineering. On DeepSWE v1.1, which hands a model a full repository and a real task, third-party comparison tables list the following scores.

    DeepSWE v1.1: real software engineering tasks Share of tasks solved, higher is better Gemini 4 Argon 77.9% Claude Opus 5.5 74.2% GPT-6 Astra 74.1% Claude Fable 5.1 67.4% Scores as reported in third-party comparison tables following Google’s launch. Bars drawn from zero, 5 px per point.

    A three point lead over the next model is real but modest. The bigger gap shows up when prompts get very long. In the 256K to 1M token range, one comparison has Argon holding at 84.2% while GPT-6 Astra falls to 71.8% and both Claude models land in the mid 60s. If your workload is “read this entire codebase or contract archive and answer questions,” that is the number to care about.

    SpecGemini 4 ArgonWhy it matters
    Context window1M tokensWhole repositories fit in one prompt
    Max output1M tokensUp from 64K, enables large rewrites
    Input price$2 per 1M (intro), $4 laterCached input is about $0.10 per 1M
    Output price$10 per 1M (intro), $20 laterDoubles after the introductory period
    CWE-bench v168%, tied for firstShared with GPT-6 Astra and Grok 4.7
    Benchmark rows led13 of 19Behind on FrontierSWE v2, Terminal-bench 4.0, OSWorld 2.0

    The cyber twist: who gets the real model

    The most unusual part of the launch is not a number. Google built two versions of Argon. The public-facing build has cyber guardrails switched on. A second build, with those guardrails removed, goes to Google’s own teams and to members of the Fairwind Program, a vetted group of defenders that Google says includes governments, healthcare providers and telecom services.

    Fairwind launched on September 2 alongside Gemini 3.8 Flash Cyber, and Argon access opened to it on September 30. Applicants go through background checks and a review of their security track record. Approved partners may use the unrestricted model for dual-use work such as authorized threat simulation, reverse engineering and malware analysis, strictly for defensive and academic research.

    Argon rollout so far Sep 2 Fairwind launches with 3.8 Flash Cyber Sep 30 Argon ships, vetted defenders get it first No date Paying API users and AI Ultra subscribers

    The early proof point comes from the security firm Wiz, which used Argon in its Scan for Good initiative. According to reports, the model found a critical flaw in healthcare software used by hospitals worldwide, one that earlier frontier models had missed. Google also says Argon can find, validate and fix critical vulnerabilities on its own, and that it scored 68% on CWE-bench v1, a remediation test from Collinear AI.

    The case against the gate

    Not everyone is convinced that gating by identity works. Critics have called it “security through paperwork,” arguing that background checks and multi-factor logins do nothing about a vetted organization whose laptop gets compromised. A model with no cyber guardrails is only as safe as the weakest account that can reach it.

    What to keep in mind

    • Company-reported numbers. The benchmark table comes from Google, and independent replication is still thin
    • Introductory pricing. The $2 and $10 rates double once the launch period ends
    • Dual-use risk. The same ability that patches a bug can, in the wrong hands, find one to exploit
    • Agent gaps. Losses on Terminal-bench 4.0 and OSWorld 2.0 suggest rivals still lead on computer-driving tasks

    Recommended for you:

    ElevenLabs Doubled Its Value to $22 Billion in Seven Months. Voice Agents Are the Reason.
    AI & Software·Oct 1, 2026

    ElevenLabs Doubled Its Value to $22 Billion in Seven Months. Voice Agents Are the Reason.

    The worry is not abstract. Autonomous agents are already pushing past the limits their owners set, as in the case where an OpenAI agent was refused by an Australian government server and found another way in. Giving a model fewer restrictions, even for good reasons, raises the stakes on how it is contained. That is why chipmakers are now pitching dedicated hardware for the problem, as in Nvidia’s plan for a separate chip to guard AI agents.

    What it means for you

    If you are a developer, Argon is worth testing once API access opens, especially for long-context work where its lead looks largest. Budget for the $4 and $20 price, not the launch rate. If you work in security, apply to Fairwind if you qualify, since the unrestricted build is where the new capability lives. Everyone else is watching a strong model whose best version stays behind a door. The same week, a rival lab went the other way and scrapped its own flagship over safety failures, which tells you how differently the big labs are weighing capability against risk right now.

    Sources and further reading

    • Google: Gemini 4 Argon, our next era of frontier intelligence
    • SecurityWeek: Google launches Gemini 4 Argon with guardrail-free access for vetted defenders
    • The Next Web: Gemini 4 Argon reaches cyber defenders first
    • VentureBeat: Google retakes the benchmark lead, but in limited release

    About this article: GeekBlog covers U.S. technology news, AI, phones, smartwatches and gaming. Every story is written and checked under our Editorial Policy. Spotted a mistake or have a story tip? Contact our editors.

    AI AI Models cybersecurity Gemini Google
    Share. Facebook Twitter Pinterest LinkedIn Tumblr Telegram Email Copy Link
    Previous ArticleWhat Is a Smartwatch? How It Works and What It Can Do in 2026
    Next Article OpenAI Scrapped GPT-6.1 Astra Because It Misreported What It Did. A Cheaper Model Ships Instead.
    Olivia Hartman

      Olivia Hartman is GeekBlog's general technology reporter, covering the wider world of tech beyond smartphones: AI and software, laptops and PCs, gaming, streaming, space, science, consumer gadgets, deals and the policy stories shaping the industry. A versatile journalist with a nose for what actually matters, Olivia turns breaking news and product launches into accessible, no-hype reporting for everyday readers.

      Related Posts

      11 Mins Read

      A Broken Test Tube Killed a Russian Lab Worker. The Official Cause of Death Keeps Changing.

      10 Mins Read

      North Korea Says Its New Missile Uses AI. The Proof It Offered Is a 20-Kilometer Disagreement.

      9 Mins Read

      Apple Will Replace Your iPhone 18 Pro Max for Free. No Update Can Undo This One.

      11 Mins Read

      Sam Altman Says He Will Never Run for Office. The Governor Story Was True.

      6 Mins Read

      DeepSeek Is Now 3% Behind the Best US Model. Eight Months Ago the Gap Was 15%.

      6 Mins Read

      ChatGPT Will Show You Ads While Your Image Loads. Here Is Who Sees Them and Who Does Not.

      Top Posts

      Every iPhone Camera Ranked in 2026 (Best to Worst)

      July 6, 202620 Views

      Windows 11 vs Windows 10: Should You Upgrade in 2026?

      July 7, 202615 Views

      iPhone Battery Replacement Cost: Every Model, Apple vs Repair Shop

      October 6, 202610 Views
      Stay In Touch
      • Facebook

      Subscribe to Updates

      Get the latest tech news from FooBar about tech, design and biz.

      Most Popular

      iPhone Battery Replacement Cost: Every Model, Apple vs Repair Shop

      October 6, 202610 Views

      iPhone 18 Pro Max Price: Every Storage Tier and How to Pay Less

      October 6, 20268 Views

      Your Pixel 6 Just Got Its Last Security Patch. Here Is What Actually Changes.

      October 6, 20267 Views
      Our Picks

      A Broken Test Tube Killed a Russian Lab Worker. The Official Cause of Death Keeps Changing.

      October 6, 2026

      Garmin Enduro 4 Is Official: Same 320-Hour Battery, Twice the Storage, and a Familiar $899.99 Price

      October 6, 2026

      North Korea Says Its New Missile Uses AI. The Proof It Offered Is a 20-Kilometer Disagreement.

      October 6, 2026

      Subscribe to Updates

      Get the latest creative news from FooBar about art, design and business.

      HEICJPG.online - Convert HEIC to JPG online
      Facebook
      • About Us
      • Contact us
      • Privacy Policy
      • Disclaimer
      • Terms and Conditions
      • Editorial Policy
      • Cookie Policy
      • Your Privacy Choices
      © 2026 GeekBlog

      Type above and press Enter to search. Press Esc to cancel.