Google has introduced Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber, three new models designed to improve efficiency, speed, and reliability for developers and customers building AI agents at scale (21/07).
The new models expand the Gemini Flash series, which focuses on balancing efficiency and quality to support agentic workflows. Google said the releases build on Gemini 3.5 Flash and provide improvements in token efficiency, latency, and performance.
Google Introduces Gemini 3.6 Flash With Improved Efficiency and Performance
Gemini 3.6 Flash is designed as a workhorse model for coding, knowledge work, and multimodal performance. Google said the model reduces output token usage by 17% compared with Gemini 3.5 Flash based on the Artificial Analysis Index, while some benchmarks, including DeepSWE by Datacurve, showed reductions of up to 65%.
The company said Gemini 3.6 Flash requires fewer reasoning steps and tool calls to complete multi-step workflows. The model is available at a lower cost, priced at $1.50 per one million input tokens and $7.50 per one million output tokens.
Google reported that Gemini 3.6 Flash delivers improved results across several use cases. The model achieved higher precision with fewer unwanted code edits and fewer execution loops in DeepSWE, scoring 49% compared with 37% for Gemini 3.5 Flash.
The model also improved performance in machine learning research, reaching 63.9% on MLE Bench compared with 49.7% for Gemini 3.5 Flash. Its computer use capability increased to 83.0% on OSWorld-Verified, compared with 78.4% for the previous model.
Google said computer use is now available as a built-in client-side tool through the Gemini API and Gemini Enterprise. The company also reported stronger knowledge work performance, with Gemini 3.6 Flash scoring 1,421 on GDPval-AA v2 compared with 1,349 for Gemini 3.5 Flash.
Customers including Hebbia and Harvey have used the model for multimodal tasks such as document parsing, chart and data analysis, and report drafting.
Gemini 3.5 Flash-Lite Designed for High-Volume AI Workflows
Google also launched Gemini 3.5 Flash-Lite, describing it as the fastest and most cost-effective model in the Gemini 3.5 series.
According to the Artificial Analysis Index, Gemini 3.5 Flash-Lite delivers 350 output tokens per second. The model is priced at $0.30 per one million input tokens and $2.50 per one million output tokens.
Google said Flash-Lite is designed for low-latency tasks and high-throughput workloads, including agentic search and document processing. Developers can configure the model to prioritize low-cost, low-latency execution or use higher thinking levels for multi-step subagent workloads.
The company said Gemini 3.5 Flash-Lite significantly outperforms Gemini 3.1 Flash-Lite across agentic workflows. It also improved performance in coding, long-context processing, and real-world task execution.
In Terminal-Bench 2.1, Gemini 3.5 Flash-Lite scored 54% compared with 31% for Gemini 3.1 Flash-Lite. In GDM-MRCR v2, it achieved 72.2% compared with 60.1%, while GDPval-AA v2 showed a score of 1,140 compared with 642.
Google also reported that Gemini 3.5 Flash-Lite outperformed Gemini 3 Flash in several evaluations, including SWE-Bench Pro with 54.2% compared with 49.6%, and OSWorld-Verified with 74.0% compared with 65.1%.
Gemini 3.5 Flash Cyber Brings AI Model to Cybersecurity
Google introduced Gemini 3.5 Flash Cyber, a specialized cybersecurity model built on Gemini 3.5 Flash and fine-tuned to identify and fix software vulnerabilities.
The model works with CodeMender, a system that uses multiple Gemini 3.5 Flash Cyber agents to produce a combined security report. Google said the system achieved competitive performance at the frontier on the CyberGym benchmark.
Google said Gemini 3.5 Flash Cyber will initially be available only to governments and trusted partners through a limited-access pilot program in CodeMender.
The company said the approach is intended to help frontline defenders identify and fix critical vulnerabilities before they can be exploited while reducing the risk of broader misuse.
Gemini 3.6 Flash and 3.5 Flash-Lite Available for Developers and Enterprises
Google said Gemini 3.6 Flash and Gemini 3.5 Flash-Lite are available starting 21 July 2026 through the Gemini API in Google AI Studio and Android Studio. Gemini 3.6 Flash is also available in Google Antigravity.
For enterprises, the models are available through the Gemini Enterprise Agent Platform, while Gemini 3.6 Flash is also available in the Gemini Enterprise application.
Google said consumers can access the models through the Gemini app, and Gemini 3.5 Flash-Lite is also rolling out in Google Search.
PHOTO: GOOGLE
This article was created with AI assistance.
We make every effort to ensure the accuracy of our content, some information may be incorrect or outdated. Please let us know of any corrections at [email protected].
Read More

Wednesday, 22-07-26
