Skip to content

Google announces Gemini 4 Argon, a frontier model for coding and cyber defense

Google's new Gemini 4 Argon targets real-world software engineering and cybersecurity defense, rolling out first to trusted defenders with introductory pricing of $2/$10 per million tokens.

By VibecodedThis 2 min read
Screenshot of the Google Gemini app interface on a phone
Image: Wikimedia Commons

Google has a new frontier model. On September 30, the company announced Gemini 4 Argon, which it describes as built for real-world software engineering, enterprise knowledge work in fields like legal and finance, and cybersecurity defense. The announcement came in a post on The Keyword by Koray Kavukcuoglu, SVP of Google DeepMind and the company's Chief AI Architect.

The rollout is staged. Argon goes first to trusted cyber defenders through Google's Fairwind Program, the limited-access program Google launched in early September for governments and partners, which now counts more than 650 participating organizations. Kavukcuoglu said releasing capabilities at this level safely requires a phased approach, and that Google is participating in the U.S. government's voluntary pre-release access process while it expands availability.

Pricing is introductory: $2 per million input tokens and $10 per million output tokens, with cached input at 95% off. A footnote in the announcement says the price rises to $4 and $20 per million tokens once the introductory period ends.

What Google says it can do

The coding claims are the headline. Google reports Argon sets a new state of the art on DeepSWE v1.1 at 77.9%, a benchmark for real-world, long-horizon software engineering tasks. It also says the model leads the Vals Index, which weights finance, coding, legal, and tax work by their contribution to U.S. GDP, and ranks first on AutomationBench at 51.3%. The output token limit jumps to 1 million, up from 64K, which Google calls industry-leading.

Google is already using Argon internally, and the examples it published are specific. Quantum computing researchers used it to optimize the spacetime resources of subroutines that bottleneck important applications, beating a published baseline by 40% in minutes. Teams of Argon agents analyzed fleet-wide profiling telemetry and applied memory optimizations across Google's data centers, freeing more than 300 TiB of memory, with an estimated 500 TiB to 1 PiB in total savings. And Argon agents are migrating C/C++ codebases to Rust across the company, from core libraries like re2 and libgav1 up to the 800K-line Fuchsia OS Zircon kernel, with rewrites going through automated and manual auditing before production.

The cyber defense pitch

Argon was trained to autonomously find, validate, and patch critical software vulnerabilities. Google says it will release the model without cyber guardrails to trusted defenders and its own internal teams. Wiz is already using it through its Scan for Good initiative, and Google says the model uncovered a critical vulnerability exposing sensitive personal information across healthcare software used by hospitals worldwide, a risk previous frontier models had missed. On CWE-bench v1, Google reports Argon ties for first at 68%.

On safety, Google lists strengthened frontier safeguards in four areas: misuse defense under its Frontier Safety Framework, monitoring of internal activations for misuse, chain-of-thought monitoring with execution stops, and hardened sandboxed environments. It also describes Argon as its most resilient model yet against indirect prompt injection, citing leading robustness on Gray Swan's benchmark.

Broad availability starts with paid API customers and Google AI Ultra subscribers, "as soon as possible." All benchmark figures are Google's own; independent verification will have to wait for wider access.