AI News Feed
Market watch
Companies

Google Releases Gemini 4 Argon, Claiming Benchmark Gains but Facing Doubts

Google launched Gemini 4 Argon, claiming top scores in 13 of 18 benchmarks, but access is limited amid internal doubts.

The name Argon comes from an internal codename kept for the official release, ifanr reported. Google did not explain the name. The element argon is an inert gas, a contrast with Chrome. The report said users should not rush to try it in the Gemini app because access remains limited.

Official benchmarks show Gemini 4 Argon leading 13 of 18 tests. Its largest advantages are in knowledge work. On Vals Finance Agent v2 and Harvey's Legal Agent Benchmark, which involve real financial analysis and real legal work, the model leads other models by what ifanr described as two tiers. Long-context performance also remains a Gemini strength. On GraphWalks, which tests use of 256k to 1M token contexts, Gemini 4 leads other flagship models by more than 10 percent.

A new feature helps explain that lead. Gemini 4 Argon raises the output limit from 64,000 tokens in the previous generation to 1 million tokens, ifanr reported. That could allow it to produce roughly 700,000 to 1 million Chinese characters in one run, depending on willingness and cost. Google's explanation, as reported, is that when a model has enough space for deep thought and can generate hundreds of thousands of tokens in a single reasoning trace, it can bring new depth to reasoning and solve difficult problems in one pass.

Coding results are mixed. On DeepSWE v1.1, a common real-world long-horizon software construction test, Gemini 4 scored 77.9 percent, nearly 4 percentage points ahead of GPT-6 Astra and Claude Opus 5.5. But it ranked last on FrontierSWE v2, which requires building projects from scratch, and on Terminal-Bench 4.0, which tests complex work in a Linux terminal. Ifanr summarized the model's profile as knowledge and writing first, coding second, terminal performance third.

Third-party results partially support that view. In Arena.ai, where blind votes determine rankings, Gemini 4 ranked first in Text Arena after high scores in difficult prompts, instruction following and creative writing. It ranked only eighth in Code Arena: WebDev and Agent Arena, which place more weight on engineering ability. The report argued that a Gemini model with weaker coding could still be valuable for writing, especially as some frontier models optimized for agentic coding have developed awkward formatting and obscure jargon.

Hallucination rates also fell. In Artificial Analysis testing cited by ifanr, Gemini 4's hallucination rate was 15 percent, the lowest among frontier models. That means it is more likely to say it does not know when facing uncertainty rather than invent an answer.

Pricing may be the most important issue for ordinary users. Google's API price for Gemini 4 Argon is $2 per million input tokens and $10 per million output tokens, with cached hits at 5 percent, or $0.1, ifanr reported. Google said the price will double after an introductory period, but ifanr said such discounts often last until the next generation. At the introductory price, Gemini 4 costs one-fifth as much as GPT-6 Astra and Fable 5.1 and is roughly level with GPT-6.1 Sol and Sonnet 5.5, making it the best intelligence-for-price model if its capabilities match the claims.

The release also ends an awkward period for Google's AI Pro subscription, which for nearly half a year advertised access to the 3 Pro model even as Flash models surpassed Pro after the 3.5 Flash era, according to the report. Many subscribers, ifanr said, were effectively paying for the 5TB Google Drive benefit. Google can now advertise its membership more directly.

Still, the report urged caution. It said large models are ultimately a manufacturing business, like phones and chips, and a long-delayed model is unlikely to suddenly surpass rivals. It pointed to a gap between benchmark scores and real-world performance for Gemini 3.8 Flash as a sign that Google may have done special optimizations for benchmarks. Almost at the same time as Gemini 4's announcement, Bloomberg reported that Google employees were skeptical of its actual performance, citing people familiar with the matter who said it struggled with certain coding tasks. Google quickly pushed back, and some employees said Gemini 4 was good, according to ifanr.

The final version that reaches users is therefore unlikely to be as strong as the benchmark table suggests, especially as ChatGPT and Claude may iterate again before Gemini 4 is widely available, ifanr said. Even so, the report said the release shows Google is back in the race. Its ceiling may not be high, but its floor is likely secure, and the next task is to catch up step by step as in any manufacturing business.