Google Gemini 4 Argon closes the gap with OpenAI and Anthropic but doesn't take a clear lead
by Matthias Bastian · The DecoderGoogle Gemini 4 Argon closes the gap with OpenAI and Anthropic but doesn't take a clear lead
Matthias Bastian View the LinkedIn Profile of Matthias Bastian
Oct 1, 2026
Key Points
- Google has unveiled Gemini 4 Argon, its latest flagship model that closes the gap with frontier models from OpenAI and Anthropic, beating some of them on key benchmarks.
- A new feature: Argon supports up to one million output tokens, letting it work through complex reasoning in a single pass without timing out.
- Google is rolling out Argon in stages, starting at an introductory price of $2 per million input tokens and $10 per million output tokens, with regular pricing rising to $4 and $20 later. Cached inputs are discounted by 95 percent.
Ask about this article… Search
Google unveiled Gemini 4 Argon, its new frontier model that closes the gap with rivals from OpenAI and Anthropic, beating some of them on key benchmarks. While it may not clearly lead the pack, it is relatively cheap for a frontier model, at least at the introductory price.
Argon is Google's first frontier model in more than seven months, following Gemini 3.1 Pro. It puts the ad giant back among the top three AI labs, though Anthropic likely still holds the lead. After a difficult and drawn-out development period that saw the already-announced Gemini 3.5 frontier model skipped entirely, Google is back in the race.
Most users will have to wait
Argon is initially going to a group of "trusted cyber defenders" as part of the Fairwind program. They and Google's internal teams will get the model without cyber guardrails. Google justifies the gradual rollout with a "phased approach" that AI capabilities at this level require. The company is also taking part in the US government's voluntary program that gives agencies access to new models before public release.
Feedback from early testers will feed into the model's safety mechanisms. Only after that does Google plan to open Argon up to developers, businesses, and consumers, starting with paying API customers and Google AI Ultra subscribers. The company hasn't given a date, saying only "as soon as possible."
Pricing is already set, at least as an introductory rate: $2 per million input tokens and $10 per million output tokens. Cached input tokens cost 95 percent less, working out to about 10 cents per million. Gemini 3.8 Flash had a 90 percent cache discount. That puts Google well below other frontier models on raw token price, though not on token consumption (see below).
| Price per million tokens | Gemini 4 Argon (promotional) | Gemini 4 Argon (regular) | GPT-6 Astra | Claude Fable 5.1 | Claude Opus 5.5 |
|---|---|---|---|---|---|
| Input Token | 2 $ | 4 $ | 10 $ | 10 $ | $4 |
| Dispensing tokens | $10 | $20 | $50 | $50 | $20 |
| Cache read operations | $0.10* | $0.20* | $1 | $0.25 | $0.20 |
| Cache write operations | N/A | N/A | $12.50 | $12.50 (5 min.) / $20 (1 hr.) | $5 (5 min.) / $8 (1 hr.) |
*Google doesn't state this figure explicitly but says the cache is 95 percent cheaper than the regular input token price.
Google also raised the output limit from 64,000 to one million tokens, calling it an industry first. The idea is that if the model can generate hundreds of thousands of tokens in a single trajectory, it can think through hard problems more thoroughly and solve them in one pass. To support this, Google is adding a new "Long Decode Continuation" feature to the Gemini API. It pauses long responses and resumes them through follow-up requests so reasoning doesn't hit a timeout.
The input context window stays at one million tokens. Argon accepts text, images, video, and audio as input but only outputs text.
Independent tests show Argon matching GPT-6 Astra but burning more tokens
Artificial Analysis provides an early independent assessment. At its highest available reasoning level, "High," Gemini 4 Argon scores 53 points on the Artificial Analysis Intelligence Index. That ties it with OpenAI's GPT-6 Astra (max) and Claude Fable 5.1, and puts it one point ahead of GPT-6.1 Sol (max).
Anthropic's models still lead. Claude Opus 5.5 sits at 58 points and Claude Sonnet 5.5 at 56. Compared to Google's last frontier model, Gemini 3.1 Pro Preview, Argon jumped 23 points. "High" is typically the top reasoning tier for Gemini models, though a special "Deep Think" mode is sometimes supported as well.
At the current promo price, one Intelligence Index task costs $1.99. That's 60 percent of GPT-6 Astra's cost ($3.26) but 2.7 times more expensive than GPT-6.1 Sol. Once the discount ends, the cost rises to $3.98, about 20 percent above GPT-6 Astra. The price advantage comes from lower token rates, not from efficiency. Argon uses an average of 62,000 output tokens per task, while GPT-6 Astra needs only 27,000.
Argon also made significant gains on agentic tasks, which according to Artificial Analysis have been a weak spot for Gemini models. On AutomationBench-AA, the Artificial Analysis variant, it takes first place at 77.5 percent, six points ahead of Claude Sonnet 5.5 (max). On Terminal Bench 4, it hits 57 percent, a 53-point jump over Gemini 3.1 Pro Preview. That still leaves it behind Claude Sonnet 5.5 (64 percent), Claude Opus 5.5 (60 percent), and GPT-6 Astra (59 percent).
Artificial Analysis also highlights Argon's low hallucination rate. On AA-Omniscience, a benchmark that tests factual knowledge and honest handling of knowledge gaps, Argon's hallucination rate is 15 percent. GPT-6 Astra (max) comes in at 51 percent and GPT-6.1 Sol (max) at 54 percent. Argon is far more likely to admit it doesn't know an answer rather than guess wrong. Its accuracy, however, reaches only 50 percent, five points below Gemini 3.1 Pro Preview and 13 points below GPT-6 Astra (max, 63 percent). On the benchmark's overall score, Argon lands at 42 points, roughly even with GPT-6 Astra (43) and GPT-6.1 Sol (42).
Google's own benchmark results paint a rosier picture. Argon leads in most of those benchmarks, sometimes by wide margins.
Argon also leads the Vals Index. At 68.9 percent, it takes first place according to Vals AI, making it the first Gemini model to top the index. Argon finishes in the top five on 20 of 22 tested benchmarks, with particular strength in finance, law, coding, and security. In this test, though, the mid-tier model Sonnet 5.5 also outranks Anthropic's top model Opus 5.5, so take it with a grain of salt.
As always, AI models have to prove themselves in real-world use, and performance depends not just on the model itself but also on the software wrapped around it. That's Google's weak spot right now: compared to Claude Cowork and ChatGPT Work, the Gemini app still lags behind.
Argon tops the human preference rankings for text
On Arena.ai, where humans rate model outputs in head-to-head comparisons, Argon performs well. In the Text Arena, Gemini 4 Argon (High) takes first place with 1,525 points, 20 points ahead of Claude Opus 4.6 (High) in second. That makes it a strong contender for writing tasks, especially after a long drought of competitive writing models and Opus 5.5 still not matching Opus 4.6 in human preference rankings. Google's previous model, Gemini 3.8 Flash (High), had been in eleventh place.
According to Arena, Argon leads in coding, hard prompts, instruction following, longer queries, and creative writing. It also ranks first across all evaluated professional fields, as well as for queries in English, Chinese, Russian, and non-English queries overall.
Web development results are more modest. In Code Arena: WebDev, Argon scores 1,679 points and lands in eighth place. That's a 96-point improvement over Gemini 3.8 Flash (High) and a jump from 29th, but it doesn't crack the top spots.
On price-to-performance, Arena puts Argon ahead of the field. At a blended rate of $8 per million tokens, Argon shifts the Pareto frontier of the Text Arena and is currently the most cost-efficient model in the ranking.
AI News Without the Hype – Curated by Humans
Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.
Subscribe now
Source: Artificial Analysis | Google Deepmind | Vals AI / Index | Arena / Bewertung