Gemini 4 puts Google back in the frontier AI race

•

Sep 30, 2026

•

8:39pm UTC

Copy link
Share on X
Share on LinkedIn
Share on Instagram
Share via Facebook

As its two biggest rivals continue to one-up each other with new frontier models, Google is making its biggest move of 2026.

On Wednesday, the search giant unveiled Gemini 4 Argon, its latest flagship frontier model. The company said that Argon offers frontier performance in a number of complex workflows, including in software engineering, knowledge work, cybersecurity, and domains such as legal and finance.

Google will begin rolling out the models specifically to cyber defenders in the Fairwind program, its restricted-access AI cyber program that it launched in early September. The company is also currently going through the US government's voluntary pre-release model checks before opening up the model to the general public.

Google noted that Argon surpassed its previous generation, Gemini 3.8 Flash Cyber, in cybersecurity tasks such as real world-vulnerability discovery and penetration testing.

The company said that Argon sets a new state of the art score on DeepSWE v1.1, the benchmark testing performance in real-world long-horizon software engineering tasks, sweeping OpenAI's Astra and Anthropic's Claude Fable 5.1 and Opus 5.5.

  • Google's Argon also outperforms these competitors in benchmarks for knowledge work, long-context tasks, and computer use.
  • Specifically, for knowledge work, the company said that Argon leads on the Vals Index, which measures impact in finance, coding, legal, and tax work, and offers state-of-the-art performance in visual understanding tasks, such as chart, document and video analysis, measuring by the LVBench for multimodal understanding.
  • However, Argon is still beat by Astra in FrontierSWE v2, which evaluates agents on complex, multi-hour technical tasks and Terminal-Bench Science for scientific workflows. Opus 5.5 also beats Argon on PostTrainBench for machine learning engineering, and Terminal-Bench 4.0 for tasks within sandboxed command-line terminal environments

In its announcement, Google noted several ways in which Argon is "fundamentally changing" its own workflows, including helping its quantum researchers optimize algorithms, improving memory efficiency, and handling large-scale codebase migrations.

The model has an output limit of 1 million tokens, up from the previous 64,000 tokens. A Google representative told The Deep View that Argon will launch at an introductory price of $2 per million input tokens and $10 per million output tokens, with cached input tokens priced at 95% off input token price. That price puts it in line with both Anthropic's Claude Sonnet 5.5 and OpenAI's GPT-6.1 Sol, which run at the same cost of $2 per million input tokens and $10 per million output tokens. However, after the introductory period, the price of Argon will double to $4 per million input tokens and $20 per million output tokens.

Google noted that the roll out will begin with paid API customers and Google AI Ultra subscribers.

Our Deeper View

Google could not have picked a more heated time to reenter the high-end frontier model race. In recent months, the company's contributions to the frontier landscape have largely focused on speed and efficiency. For instance, its early September release of Gemini 3.8 Flash cost a fraction of what its frontier competitors were charging at the time. But in recent weeks, with both Anthropic and OpenAI homing in on token efficiency as well, Google's Argon may not be able to compete on just state-of-the-art performance alone, especially as model labs continue to leapfrog each other in capability and efficiency week after week. What a model costs is quickly starting to matter as much as what it can do.