Google Opens Gemini 4 Argon to Vetted Cybersecurity Defenders First
Google began rolling out Gemini 4 Argon, a frontier model it says leads Anthropic and OpenAI on most of its benchmarks, to vetted cybersecurity defenders, with general developer access still undated.
Outside Google's own teams, only members of its Fairwind Program can use Argon for now. The program opened Sept. 3 with the smaller Gemini 3.8 Flash Cyber model and has since signed up more than 650 organizations, including CrowdStrike Holdings Inc. and Palo Alto Networks Inc.
Koray Kavukcuoglu, Google's chief AI architect, wrote in the announcement that releasing capabilities at this level "requires a phased approach." Google is taking part in the U.S. government's voluntary process for pre-release model access and says it is still hardening safeguards against misuse for cyberattacks or weapons development. By Google's account, Argon withstands indirect prompt injection better than any model it has shipped. Separate monitors watch the model's chain of thought and actions and can stop it if it goes beyond what a user intended.
On DeepSWE v1.1, which measures long-horizon software engineering tasks, Argon scored 77.9%. Anthropic's Claude Opus 5.5, released the previous week, managed 74.2% on Google's chart, a tenth of a point ahead of OpenAI's GPT-6 Astra. AutomationBench, which measures end-to-end business work, showed a wider gap of 51.3% to 42.5% over Opus 5.5. SiliconANGLE reported that VentureBeat counted 12 of the 18 benchmarks in Google's charts where Argon led outright. Terminal-Bench 4.0 still belongs to Opus 5.5, and FrontierSWE v2 is one of the few Astra kept.
Google built Argon for the kind of long-horizon work DeepSWE measures and set the model's output limit at 1 million tokens to match, up from 64,000 for earlier Gemini models.
Thousands of Google employees are already using Argon, and teams at the company have put its agents to work on large engineering jobs. One effort is moving C and C++ code to Rust, in projects that range from core libraries of tens of thousands of lines to the more than 800,000-line Zircon kernel in Google's Fuchsia operating system. On libgav1, Google's open-source video decoder, the agents started from an existing Rust port and rewrote 32,000 lines of its speed-critical code as safe Rust the compiler could optimize on its own. The decoder now runs 2.7 times faster than that earlier port, with no change to its video output. Google engineers are also using Argon agents to hunt for wasted memory, and a sweep of fleetwide profiling data has so far freed more than 300 tebibytes across the company's data centers.
Fairwind members and Google's internal teams get a version of Argon with the cyber guardrails removed. The model can find and fix software vulnerabilities on its own, and on CWE-bench v1, a remediation test, it tied GPT-6 Astra at 68%.
Wiz Inc., the Google-owned cloud security company, has used Argon in its free Scan for Good program, which looks for exposures in essential public infrastructure. Google said the model found a critical flaw exposing personal information in health care software used by hospitals worldwide that earlier frontier models had not caught.
Google has not given a date for paying developers and Google AI Ultra subscribers, who are next in line. Argon will cost $2 per million input tokens and $10 per million output tokens at launch, and cached input is priced 95% lower. Anthropic charges $4 and $20 for Opus 5.5, and Argon moves to those rates when the launch pricing ends.