Skip to content
AI News and Updates

Gemini 3.6 Flash Is Out, and Google Just Confirmed Gemini 4 Pretraining

Sofia Almeida
Jul 23, 2026  /  7 min read
Green circuit board representing the chips and infrastructure behind Google Gemini AI models
Photo by BotheredByBees (CC BY), via Openverse.

Google shipped three new Gemini models on July 21, 2026 — Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and a security-focused Gemini 3.5 Flash Cyber — and, almost as an aside, confirmed it has already started pretraining Gemini 4. I’ve been tracking Google’s Flash releases since Gemini 2.5 Pro Deep Think landed, and this is the first time I’ve seen the company publicly overlap a shipping “budget” model with a confirmed next-generation flagship still in training. That’s not normal cadence. It’s a signal.

If you use the Gemini app, build on the API, or just want to know whether it’s worth switching your workloads over, here’s what actually changed and why the Gemini 4 mention matters more than the specs table Google put out.

What Google Actually Released on July 21

Three models landed the same day, each aimed at a different job:

  • Gemini 3.6 Flash — the new default mid-tier model, built for coding, knowledge work, and everyday multimodal tasks.
  • Gemini 3.5 Flash-Lite — a stripped-down, high-throughput model for agentic search and document processing where speed matters more than depth.
  • Gemini 3.5 Flash Cyber — a specialized security model that hunts for and helps patch software vulnerabilities, folded into Google’s CodeMender tool.

Both 3.6 Flash and 3.5 Flash-Lite are live now in the Gemini app, Search, and the developer API. Flash Cyber is not generally available — it’s restricted to a limited pilot for governments and “trusted partners,” which tells you Google isn’t ready (or willing) to hand a vulnerability-hunting model to the general public just yet.

The timing is notable too. This launch came days after the European Commission ordered Google to open Android to rival AI assistants, so Google is simultaneously defending Gemini’s ecosystem advantage in Europe and racing to keep its model lineup ahead of Anthropic and OpenAI everywhere else.

Gemini 3.6 Flash: Cheaper and Noticeably Better at Coding

The headline number is efficiency: Gemini 3.6 Flash uses 17% fewer output tokens than 3.5 Flash for comparable answers, which is why Google could drop output pricing from $9.00 to $7.50 per million tokens while input stayed around $1.50 per million. On paper that’s a real cost cut for anyone running high-volume API traffic, not a marketing rounding error.

The benchmark jumps are bigger than I expected for a Flash-tier refresh:

  • DeepSWE coding benchmark: 49%, up from 37% on 3.5 Flash
  • MLE-Bench: 63.9%, up from 49.7%
  • GDPval-AA (real-world knowledge work): 1421, up from 1349
  • OSWorld-Verified (computer-use tasks): 83%, up from 78.4%

Google also pushed the model’s knowledge cutoff forward to March 2026, a meaningful jump from the January 2025 cutoff on 3.5 Flash. For anyone using Gemini for current-events questions or recent library documentation, that alone cuts down on a lot of “I don’t have information past my training data” responses.

Gemini 3.5 Flash-Lite: Built for Volume, Not Depth

Flash-Lite is the model you reach for when you’re firing off thousands of small requests and latency is the bottleneck — agentic search loops, document extraction pipelines, that kind of thing. Pricing lands at $0.30 per million input tokens and $2.50 per million output tokens, and Google says it runs at roughly 350 output tokens per second.

The quality gains over its March predecessor are bigger than the “Lite” name suggests: Terminal-Bench 2.1 nearly doubled from 31% to 54%, long-context performance on GDM-MRCR v2 climbed from 60.1% to 72.2%, and real-world execution scores on GDPval-AA v2 jumped from 642 to 1140. If your workload was avoiding Flash-Lite because the older version felt too thin, it’s worth another look.

Gemini 3.5 Flash Cyber: A Vulnerability Hunter Only Governments Get

This is the model I find most interesting, and not for benchmark reasons. Flash Cyber is purpose-built to discover, validate, and help patch software vulnerabilities, and Google ran it against the V8 JavaScript engine as a test case: 55 unique confirmed issues found, versus 47 for regular Gemini 3.5 Flash and 36 for Anthropic’s Opus 4.6.

That’s a genuinely strong result — beating a rival lab’s flagship at a security-specific task with what’s technically a mid-tier model. But Google isn’t opening it up. It’s routed exclusively through CodeMender for governments and trusted partners, which is the same pattern we’ve seen with other dual-use AI security tooling: build the capability, then gate access instead of releasing weights or broad API access. That’s a defensible call given how directly this kind of model could be repurposed for offense instead of defense.

The Real Story: Gemini 4 Is Already Training

Buried in the announcement, Google confirmed it has begun what it called its most ambitious pretraining run yet, for Gemini 4. No release window, no benchmark teasers, nothing beyond the confirmation that the run has started.

Here’s why I think that’s the bigger news than the three Flash models combined: Gemini 3.5 Pro — the actual flagship successor people have been waiting on — has now missed its ship target multiple times and still isn’t out. Google shipping a round of Flash refreshes while quietly starting pretraining on the model after the delayed one reads like a company managing a gap, not celebrating a launch. Flash releases keep developers and pricing-sensitive customers happy while the flagship pipeline sorts itself out.

I’d treat any specific Gemini 4 release date you see elsewhere online as speculation until it comes from Google directly. As of this writing, the company has only confirmed the pretraining run exists — nothing about scale, timeline, or capability targets. (9to5Google has the full breakdown of the launch if you want the raw benchmark tables.)

Quick Reference: What Changed and What It Costs

ModelBest ForInput / Output Price (per 1M tokens)Availability
Gemini 3.6 FlashCoding, knowledge work, multimodal$1.50 / $7.50Live now (app, Search, API)
Gemini 3.5 Flash-LiteHigh-throughput agentic search, document processing$0.30 / $2.50Live now (app, Search, API)
Gemini 3.5 Flash CyberVulnerability discovery and patchingNot publicly pricedLimited pilot (governments, trusted partners)
Gemini 3.5 ProFlagship reasoningN/AStill not shipped; multiple missed targets
Gemini 4Next-gen flagshipN/APretraining confirmed underway; no date given

Should You Switch to Gemini 3.6 Flash?

If you’re already on 3.5 Flash for coding or agentic tasks, yes — the DeepSWE and OSWorld gains are large enough to notice in practice, and the token efficiency means you’re likely to pay less for the upgrade, not more. If you’re running high-volume, low-complexity workloads, Flash-Lite’s jump on Terminal-Bench and long-context handling makes it worth re-benchmarking against whatever you’re using now, especially if you passed on the March version. Flash Cyber isn’t something most of us can touch yet, so there’s no decision to make there beyond watching whether Google widens access later.

Frequently Asked Questions

Is Gemini 3.6 Flash available to everyone right now?

Yes. Gemini 3.6 Flash and Gemini 3.5 Flash-Lite are both live as of July 21, 2026, in the Gemini app, Google Search, and the developer API. Gemini 3.5 Flash Cyber is not — it’s limited to a pilot program for governments and trusted partners.

What happened to Gemini 3.5 Pro?

It still hasn’t shipped. Google has missed its target for Gemini 3.5 Pro multiple times, and this announcement didn’t include a new release date for it. The company’s attention on this update was on the Flash tier and on confirming Gemini 4 pretraining.

When will Gemini 4 be released?

Google hasn’t said. The company confirmed it has started what it describes as its most ambitious pretraining run yet, but gave no timeline, benchmark targets, or capability details. Treat any specific date you see for Gemini 4 elsewhere as unconfirmed speculation.

Is Gemini 3.6 Flash actually cheaper than 3.5 Flash?

For output-heavy workloads, yes. Output pricing dropped from $9.00 to $7.50 per million tokens, and the model uses 17% fewer output tokens to produce comparable answers — so the effective cost drop for typical usage is larger than the sticker price change alone suggests.

Google’s pattern here — refresh the cheap tier, gate the security model, and quietly start training the next flagship — is worth watching over the next few months. If Gemini 3.5 Pro keeps slipping while Gemini 4 pretraining continues, don’t be surprised if Google skips straight from a delayed 3.5 Pro to a Gemini 4 launch instead.

Written by
Sofia Almeida

Sofia follows emerging technology, from AI and VR to IoT and blockchain, and translates the hype into plain language. She cares about what these tools mean for everyday users, not just the headlines.

Up Next