Google Releases Three Gemini Models: 3.6 Flash Becomes Mainstay, but 3.5 Pro Still Delayed

Google has unveiled three new Gemini models: Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber. The Gemini 3.6 Flash emerges as the new primary model, while the release timeline for Gemini 3.5 Pro remains undetermined.

July 21, 2026 – Google released three models: Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber. Gemini 3.6 Flash becomes the new flagship model, while the release date for Gemini 3.5 Pro remains uncertain.

Fact Reconstruction

Google DeepMind launched the above three models on July 21, 2026. Gemini 3.6 Flash is optimized based on feedback from Gemini 3.5 Flash, with improvements in programming, knowledge processing, and multimodal capabilities. The Artificial Analysis Index test shows an average 17% reduction in output tokens. Pricing is adjusted to $1.5 per million input tokens and $7.5 per million output tokens. Gemini 3.5 Flash-Lite targets low-latency, high-throughput scenarios, achieving an output speed of 350 tokens per second, with pricing reduced to $0.3 per million input tokens and $2.5 per million output tokens. Gemini 3.5 Flash Cyber is fine-tuned for vulnerability discovery, verification, and automated remediation, initially piloted with government and trusted partners via the CodeMender platform.

Mechanism Breakdown

This release revolves around the operational needs of production-grade AI agents. When developers build large-scale agent workflows, token consumption, response latency, and tool invocation frequency directly determine costs. Gemini 3.6 Flash reduces overall inference steps by minimizing invalid code modifications and repetitive execution loops. Gemini 3.5 Flash-Lite supports different levels of thinking depth configuration, allowing developers to balance low latency and complex reasoning. It also includes built-in Computer Use capabilities, enabling cross-browser, mobile, and desktop task execution. Gemini 3.5 Flash Cyber strengthens security protection in CBRN and cyberattack domains, reducing jailbreak attack risks while minimizing false rejections in normal scenarios.

Industry Impact

In terms of competitive landscape, Google uses the Flash series to fill the gap between high performance and low cost, forcing other vendors to follow suit on agent efficiency metrics. For developers, Gemini 3.6 Flash and Gemini 3.5 Flash-Lite are now available on the Gemini API, Google AI Studio, and Android Studio, ready for building production applications. For enterprise users, the Gemini Enterprise Agent Platform is open for access, reducing the cost of document processing, search, and code maintenance scenarios. For upstream and downstream hardware and cloud service providers, low-token-consumption models may reduce dependence on high-end GPUs, but cybersecurity-specific models may increase procurement of trusted computing environments.

Comparisons and Precedents

Unlike previous competition centered on model capability ceilings, this update emphasizes operational efficiency. Gemini 3.5 Flash-Lite outperforms Gemini 3 Flash in some agent and programming benchmarks, indicating that lightweight models already have substitution potential in specific scenarios. Gemini 3.5 Pro was originally scheduled for June but was postponed due to internal testing failing to meet targets, particularly in code generation capabilities. It is still in partner testing.

Strategic Assessment

Google has initiated its largest pre-training effort to date for Gemini 4. The next phase is most likely to see the official release of Gemini 3.5 Pro alongside early benchmark comparisons with Gemini 4.