Google’s Gemini 4 Argon narrows the frontier gap, but does not clearly lead

Google’s Gemini 4 Argon matches OpenAI’s GPT-6 Astra on one index, trails Anthropic’s top models, and launches first to selected testers.

AI News

Google has unveiled Gemini 4 Argon, a new flagship model that brings the company back into direct competition with OpenAI and Anthropic after more than seven months without a new frontier release. Early testing suggests Argon is substantially stronger than Google’s previous model and competitive with OpenAI’s GPT-6 Astra, but Anthropic’s leading systems still appear to hold the overall advantage.

The release is notable less for a single benchmark victory than for the trade-offs it exposes. Argon offers a one-million-token output limit and unusually low introductory token prices, while independent testing indicates that it consumes far more tokens than some rivals to complete the same tasks. Google is also limiting access while it gathers safety feedback, meaning developers cannot yet evaluate the model through a broadly available API.

A staged release with unusually long outputs

According to The Decoder, Google is first providing Argon to selected “trusted cyber defenders” through its Fairwind program, as well as to internal teams. The initial version reportedly operates without cyber guardrails in that testing environment. Google is also participating in a U.S. government voluntary program that gives agencies early access to new models.

The company plans to use tester feedback to refine Argon’s safety mechanisms before expanding access to developers, businesses, and consumers. Paid API customers and Google AI Ultra subscribers are expected to receive access next, although Google has not provided a firm public date.

The model’s headline capability is its one-million-token output limit, up from 64,000 tokens on the previous generation. Argon retains a one-million-token context window and accepts text, images, video, and audio, but produces text only. Google is adding a Gemini API feature called Long Decode Continuation, which can pause a long response and resume it through follow-up requests instead of allowing a lengthy reasoning process to time out.

That design targets complex research, coding, and agentic workflows in which a model may need to sustain a long chain of reasoning. It also creates a practical question for buyers: whether a larger output ceiling produces better results at an acceptable cost, rather than simply generating more text.

Independent tests show progress, not dominance

The strongest available performance evidence comes from Artificial Analysis, as reported by The Decoder. At its highest reported reasoning setting, Argon scored 53 on the Artificial Analysis Intelligence Index. That tied OpenAI’s GPT-6 Astra and Anthropic’s Claude Fable 5.1, and placed it one point ahead of GPT-6.1 Sol. Anthropic’s Claude Opus 5.5 scored 58, while Claude Sonnet 5.5 scored 56.

The result represents a major improvement over Google’s Gemini 3.1 Pro Preview, which Argon reportedly surpassed by 23 points. However, a tie with one OpenAI model is not the same as taking a clear lead across the frontier. The Decoder’s assessment places Anthropic ahead on the index, while the benchmark results vary significantly by task.

Argon performed especially well on some agentic evaluations. Artificial Analysis reported a 77.5 percent score on AutomationBench-AA, six points above Claude Sonnet 5.5. On Terminal Bench 4, Argon reached 57 percent, a substantial gain over Gemini 3.1 Pro Preview but below Claude Sonnet 5.5 at 64 percent, Claude Opus 5.5 at 60 percent, and GPT-6 Astra at 59 percent.

The model also recorded a lower hallucination rate than the compared systems on AA-Omniscience: 15 percent, versus 51 percent for GPT-6 Astra and 54 percent for GPT-6.1 Sol. Yet its accuracy on that test was 50 percent, below Gemini 3.1 Pro Preview and GPT-6 Astra. That distinction matters for enterprise applications: refusing to guess can improve reliability, but it does not automatically mean the system knows more.

Google’s own benchmark results reportedly show Argon leading in many categories, while Vals AI placed it first with a 68.9 percent score. Those claims are vendor-linked or based on selected external evaluations and should not be treated as a universal ranking. The Decoder also noted that the Vals results placed Anthropic’s mid-tier Sonnet 5.5 above its flagship Opus 5.5, a reason to interpret the ranking cautiously.

Low token prices meet high token consumption

Google is introducing Argon at $2 per million input tokens and $10 per million output tokens. The regular prices are expected to rise to $4 and $20, respectively. Cached inputs receive a 95 percent discount, which would put the introductory cached-input rate at about 10 cents per million tokens based on the figures reported by The Decoder.

Those rates appear low for a frontier model, but token price alone does not determine application cost. Artificial Analysis reported that Argon used an average of 62,000 output tokens per Intelligence Index task, compared with 27,000 for GPT-6 Astra. As a result, the estimated cost of an Argon task was $1.99 at the promotional rate, versus $3.26 for Astra. Once Argon’s regular price takes effect, the same calculation rises to about $3.98.

For builders, the implication is straightforward: Argon may be economical for workloads that benefit from long reasoning and can tolerate additional output, but less so when latency, token volume, or predictable spend matters. Teams will need to measure complete workflow cost rather than compare published per-token prices.

The model’s performance outside raw reasoning is also uneven. Arena.ai reportedly ranked Gemini 4 Argon first in its Text Arena, including coding, instruction following, hard prompts, and creative writing. In Code Arena’s WebDev test, however, it ranked eighth. That gap reinforces the importance of evaluating a model inside the software stack, tools, retrieval systems, and user interface where it will actually operate.

What to watch next

The immediate signal will be Google’s public API release and whether the company maintains the introductory price long enough for developers to test production workloads. Access for paying API customers and Google AI Ultra subscribers is planned, but no date has been announced.

Enterprise buyers should watch independent measurements of latency, output-token usage, refusal behavior, and tool-use reliability. The one-million-token output limit will matter only if long responses improve task completion without creating runaway costs or difficult-to-audit agent behavior.

The broader competitive signal will come from product integration. The Decoder reported that Google’s Gemini app still trails products such as Claude Cowork and ChatGPT Work, despite Argon’s strong benchmark results. If Google cannot pair the model with better agent interfaces, tools, and workflow controls, benchmark gains may have limited impact on adoption.

Creati.ai perspective

Gemini 4 Argon looks like a credible return to the frontier rather than a decisive takeover. Its progress is meaningful: it reportedly closes much of the gap with OpenAI, improves Google’s weak results on several agentic tests, and offers a compelling headline price. But Anthropic’s higher Intelligence Index scores and Argon’s heavy token consumption complicate the cost-performance story.

For AI builders, the release is best understood as another strong option that demands workload-specific testing. The winners in this round will not be determined by a single leaderboard. They will be determined by which model delivers reliable results, manageable inference costs, useful tool execution, and safe deployment once real users—not selected testers—can access it.

Ads