Google Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber launch hero image Image: Google
by Michael Joiner

Google Ships Gemini 3.6 Flash, 3.5 Flash-Lite, and a Restricted Security Model While 3.5 Pro Stays Delayed

Three new Gemini models landed on July 21, led by Gemini 3.6 Flash — cheaper, faster, and more accurate than its predecessor. Meanwhile Google confirms Gemini 4 pretraining has started, and 3.5 Pro is still stuck in partner testing.

Share

Google shipped three Gemini models on July 21: Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber. Each targets a different slice of the market, and none of them is the flagship model developers have been waiting for since I/O.

Gemini 3.6 Flash: Cheaper and Better Than 3.5 Flash

The headline model drops output pricing to $7.50 per million tokens, down from $9 for 3.5 Flash, while also generating roughly 17% fewer output tokens to produce the same results. That combination means lower costs without cutting capability.

The performance numbers are the more interesting part. On DeepSWE, a coding benchmark for agentic software engineering work, 3.6 Flash scores 49% against 37% for 3.5 Flash. On OSWorld-Verified, the computer use benchmark, it goes from 78.4% to 83%. MLE-Bench also improved from 49.7% to 63.9%. These aren’t marginal bumps.

The model keeps the 1 million token context window, 64k maximum output, native multimodal input, thinking controls, and built-in Computer Use from its predecessor. The knowledge cutoff moves from January 2025 to March 2026, which matters for any task that touches recent code, documentation, or events.

Pricing: $1.50 per million input tokens, $7.50 per million output tokens.

Gemini 3.5 Flash-Lite: Speed and Price Floor

Flash-Lite positions itself as the “high-volume, low-cost” option. It runs at 350 output tokens per second and prices at $0.30 per million input tokens and $2.50 per million output. Both figures put it near the bottom of the frontier model market.

On Terminal-Bench 2.1, a test for CLI agent tasks, Flash-Lite scores 54% against 31% for the 3.1 Flash-Lite it replaces. Google says it “significantly outperforms” the prior version in agentic workflows, which is the use case most likely to eat large amounts of tokens at scale.

Both Flash models share the 1M context window and the same full capability set.

Gemini 3.5 Flash Cyber: A Security Model on a Leash

The third model is not available to the general public. Gemini 3.5 Flash Cyber is purpose-built for vulnerability detection and remediation, integrated into Google’s CodeMender agent, and available only through a “limited-access pilot exclusively for governments and trusted partners.”

Google is being explicit about why. Dual-use risk — the same capability that finds and patches vulnerabilities can also discover them for offensive purposes — puts this in a restricted category. The approach mirrors what Anthropic has done with its security-focused model tier, which also requires approval gating.

Gemini 3.5 Pro: Still Not Out

Google announced 3.5 Pro at I/O in May with a June target. It missed. The current status is “in limited testing with partners,” and Google’s only public comment is that it will ship “when it’s ready.” There’s no new date.

This matters because 3.5 Pro was supposed to go directly at Claude Fable 5, GPT-5.6 Sol, and the other flagship-tier models that have already launched. Shipping three Flash-tier models while the flagship sits in testing doesn’t answer that question.

Gemini 4 Has Started Pretraining

Google closed the announcement with a note that pretraining for Gemini 4 has begun, describing it as “our most ambitious pretraining run yet.” Pretraining is the first major phase of building a model, so this means Gemini 4 is under active development, not just planned. No timeline or capability details were shared.

All three new models are available on the Gemini API and Vertex AI. Flash-Lite rolls out to Gemini apps over the coming weeks.

Share