Google Releases Gemini 4 Argon With Top Benchmarks, Lower Pricing and Cybersecurity Focus
Google released Gemini 4 Argon, a flagship model it says leads several benchmarks, costs less on some tasks and is initially limited to vetted cybersecurity defenders.
QbitAI reported that Argon scored 77.9% on DeepSWE v1.1, a benchmark for long-horizon software engineering tasks, above Claude Opus 5.5 at 74.2% and GPT-6 Astra at 74.1%. It tied for first with GPT-6 Astra on the cybersecurity vulnerability repair test CWE-bench v1 at 68%. The model topped the Vals Index with 68.90%, a large improvement over Gemini 3.8 Flash, and rose to fourth on the RSI index, surpassing GPT-6 Astra.
In third-party evaluations, Argon ranked first on Text Arena with 1,525 points, with advantages in code writing, difficult prompts, instruction following, long queries and creative writing. It placed eighth on Code Arena for web development, up 21 places from Gemini 3.8 Flash High but still behind the latest Claude and GPT models. On Artificial Analysis, it was roughly comparable to GPT-6 Astra and one point above GPT-6.1 Sol, with gains attributed to fewer hallucinations and stronger agent capabilities.
Google set Argon's maximum output at one million tokens. Pricing was listed at $2 per million input tokens and $10 per million output tokens. QbitAI said the lowest single-task cost could be $1.99, about half the cost of Astra, while also reporting that during a promotional period the per-task cost was about 40% lower than Astra's. The model is aimed at complex, long-running reasoning work such as software engineering, finance, legal workflows and cybersecurity defense.
Google cited internal uses. It said Argon autonomously identified and deployed memory optimization plans across its data centers, expected to free more than 300 TiB of memory once fully deployed. In another case, an Argon agent worked on a Rust port of a video decoder, repeatedly experimenting and replacing 32,000 lines of SIMD code; the final version ran 2.7 times faster than the original Rust port, according to the report.
The company emphasized defensive cybersecurity. Argon can autonomously discover, validate and patch critical software vulnerabilities in more than 20 languages, QbitAI reported. In an internal black-box penetration test by Wiz, the model analyzed a network system in real time without source code, found attack surfaces, identified vulnerabilities and generated proof-of-concept evidence. To limit abuse, Argon's first release is part of Google's Fairwind program for vetted cyber defenders. For trusted defensive teams, Google will relax some restrictions on cybersecurity tasks, letting Argon help defenders patch holes and search for attack entry points.
Availability remains limited. QbitAI said Argon is open only to selected cybersecurity teams, with other subscribers to follow gradually. Google executives Sundar Pichai, Demis Hassabis and Koray Kavukcuoglu publicly backed the model. DeepMind researcher Zirui Wang wrote in a post that "RSI is here" and joked that it was time to pursue his barista dream.
Not all internal reactions were positive. QbitAI cited a Bloomberg report saying some employees who had contact with Gemini 4 believed the model did not perform as strongly as its benchmark scores in actual programming tasks, that front-end design for web and app interfaces remained a weak point, and that there were suspicions of over-optimizing for leaderboards. Another employee familiar with development work said the internal view was that Gemini 4 had reached the frontier. Google responded that it was inaccurate to say the model performs poorly at programming.
For general users, the model is not yet available. QbitAI said whether Argon lives up to its scores will have to be observed in practice.
Editor's Summary
Google released Gemini 4 Argon with leading scores on several benchmarks, lower pricing than Astra on some tasks, and an initial rollout restricted to vetted cybersecurity defenders. The launch marks Google's return to the flagship model race, but internal skepticism and limited public access mean its real-world performance remains to be tested.