- Gemini 3.6 Flash cuts output token usage by 17% compared to 3.5 Flash while delivering measurable gains in coding, knowledge work, and agentic tasks.
- Gemini 3.5 Flash Cyber, built to find and fix security vulnerabilities, is restricted to governments and trusted partners under a limited-access pilot with no commercial release date confirmed.
Julian Goldie posted a YouTube video on July 22, 2026, catching these model releases almost as they went live. He spotted Gemini 3.6 Flash and 3.5 Flash-Lite inside Google AI Studio before the news reached most people. By the time he stopped recording, 3.6 Flash had already pushed to every account he owned.
According to Goldie, the knowledge cutoff of 3.5 Flash was January 2025, but the knowledge cutoff of 3.6 Flash went all the way to March 2026. Goldie also mentioned that there was a change in product name, where the old name was Notebook LM but it is now known as Gemini Notebook, and it can even write and execute code on synchronization from the Gemini app.
The Google DeepMind official release happened on July 21, 2026. There were three models announced. The Google blog post about the same release was written by Tulsee Doshi, the Senior Director of Product Management, Gemini team. The entire announcement had one focus area, which was efficiency, latency, and reliability for the developers using the AI agent.
Google’s Workhorse Gets Leaner Without Losing Ground
Google labeled 3.6 Flash its workhorse. Based on the Artificial Analysis Index, the model cuts output token usage by 17% compared to 3.5 Flash. Specific benchmarks push that gap higher; DeepSWE by Datacurve shows reductions of up to 65% in some workflows. Pricing reflects that efficiency push, landing at $1.50 per million input tokens and $7.50 per million output tokens, under what 3.5 Flash costs.
Coding precision moved up: DeepSWE scores went from 37% to 49%, with fewer unwanted edits dragging down those numbers. Computer use hit 83.0% on OSWorld-Verified, up from 78.4%. Knowledge work benchmarks reflected a similar movement, GDPval-AA v2 climbing from 1349 to 1421. Figma, Harvey, Hebbia, and JetBrains all appeared in the announcement as early customers validating real-world results.
Developers can reach both 3.6 Flash and 3.5 Flash-Lite right now through Google AI Studio and Android Studio. For 3.6 Flash specifically, Google Antigravity is also an access point.
3.5 Flash-Lite: Speed-First, Budget-Friendly
Gemini 3.5 Flash-Lite was built for volume. Artificial Analysis clocks it at 350 output tokens per second, and pricing sits at $0.30 per million input tokens and $2.50 per million output tokens, according to the Android Authority breakdown of the launch.
Google claims it clears not just 3.1 Flash-Lite but Gemini 3 Flash as well in agentic tasks. SWE-Bench Pro: 54.2% versus 3 Flash’s 49.6%. OSWorld-Verified: 74.0% versus 65.1%. For teams running parallel agent workloads at high volume, Flash-Lite becomes a legitimate cost-control layer below Flash.
Gemini 3.5 Flash Cyber: The Model Behind a Gate
The third release runs by different rules. Google built 3.5 Flash Cyber on top of 3.5 Flash, fine-tuned for one purpose: detecting, validating, and patching code vulnerabilities at a cost point lower than what larger models demand.
Inside CodeMender, Google’s managed code security agent, multiple 3.5 Flash Cyber instances work in tandem and produce one consolidated report. TechCrunch noted the model is already running inside Google’s internal codebases, patching issues across Android, Chrome, and YouTube. Who can use it? Governments and trusted partners only, through CodeMender, under a limited-access pilot, as confirmed by CNET’s coverage. No commercial release window has been confirmed.
Gemini 3.5 Pro: Still Not Here
What Google did not ship drew just as much attention. Gemini 3.5 Pro was previewed during the 3.5 Flash launch in May 2026, with Google saying it was already in internal use and would arrive “next month.” That month came and went.
There were earlier reports from Bloomberg about internal delays related to benchmarks for 3.5 Pro that have not yet been approved. In a message dated July 21, Google DeepMind product lead Logan Kilpatrick stated that the model is still under partner testing but hopes to get it soon; however, no date was provided. The 9to5Google report about the release corroborated Google’s statement about 3.5 Pro being “currently testing with partners” and “available as soon as it’s ready.”
Competitors kept moving. Since Google’s last Gemini Pro update in February 2026, OpenAI shipped GPT-5.5 and started rolling out GPT-5.6. Anthropic released both Claude Opus 4.8 and Claude Sonnet 5 in the same window, per TechCrunch’s reporting on the competitive context.
Gemini 4: Pre-Training Already Running
Kilpatrick confirmed Google has started pre-training for Gemini 4, calling it the team’s most ambitious run to date. No release window came with that. The signal is directional: DeepMind’s attention is already past the 3.x generation, as stated in the official Google announcement.
Why This Matters for Teams Building on Gemini?
Token efficiency at 17% may sound modest until you run it across millions of agent calls. The cost difference compounds. Teams on 3.5 Flash should run 3.6 Flash against their actual pipelines rather than taking the benchmark numbers as settled; improvements that hold in internal evals do not always survive contact with real production workloads.
Flash-Lite’s position is practical in a specific way: when throughput is the bottleneck, not intelligence, $0.30 per million input tokens gives teams a cost tier that wasn’t there before. Computer use as a native built-in tool through the Gemini API cuts an integration step that previously required extra setup.
For commercial teams, 3.5 Flash Cyber is inaccessible right now, which limits its immediate relevance. The architecture matters anyway: a specialized model embedded inside a managed agent system, rather than released as a standalone, says something about how Google plans to handle AI deployment in sensitive domains. As for 3.5 Pro, three months past the internal deadline and still without a firm date, teams that held off on architectural decisions around Google’s flagship model are still holding off.