Key takeaways
- Developers and customers building production AI agents need higher token efficiency, lower latency, and more reliable performance.
- In parallel, our team is already focusing on building the next generation of models.
- It also takes fewer reasoning steps and tool calls to accomplish multi-step workflows.
What happened
Developers and customers building production AI agents need higher token efficiency, lower latency, and more reliable performance. Our Flash series of models is built to meet the sweet spot of efficiency and quality to enable scaling agentic workflows. 5 Pro is currently testing with partners and we plan to make it broadly available as soon as it’s ready.
The model now also has computer use as a built-in tool to reliably support these agentic tasks across surfaces. 2% vs. 1%), and real-world task execution as seen in GDPval-AA v2 (1140 vs. 642). 2% vs. 0% vs. 5 and 3 Flash. 5 Flash-Lite model card. AI models have become capable of finding security vulnerabilities faster than current systems can fix them.
Why it matters
In parallel, our team is already focusing on building the next generation of models. We have started our most ambitious pre-training run yet, for Gemini 4, and are excited by the progress. 5 Flash. 6 Flash not only delivers a step up in coding and knowledge work, but it does this while meaningfully improving token efficiency. 5 Flash.
It also takes fewer reasoning steps and tool calls to accomplish multi-step workflows. 5 Flash. 6 Flash reduces the overall cost per agentic task, making agents more cost-effective to build and run. 6 Flash is shipping with enhanced Frontier Safety safeguards in the domains of Chemical, Biological, Radiological, and Nuclear (CBRN) and cyber offense misuses. These safeguards make the model substantially more resistant to jailbreaks.
At the same time, the model has been trained to minimize refusals for beneficial uses. 6 Flash model card. 5 Flash-Lite, designed for both low-latency tasks and tasks where high throughput is critical for developers workflows, like agentic search and document processing. 5 series. As measured by Artificial Analysis, it runs at 350 output tokens/s.
5 Flash-Lite offers a strong price-to-performance ratio for developers and customers running high throughput production traffic. 5 Flash-Lite enables efficient scaling for agentic systems. 1 Flash-Lite. Depending on the workload, developers can configure the model to prioritize low-latency, low-cost execution for high-volume tasks with the minimal and low thinking levels, or engage higher thinking levels to process multi-step subagent workloads.
What to watch
Tackling this growing threat requires an approach to securing software that is highly capable and efficient. Flash’s performance and efficiency makes it an ideal foundation to detect, validate, and patch code security issues at scale. 5 Flash, and fine-tuned for finding and fixing cybersecurity vulnerabilities at a lower price per token than larger models. 5 Flash Cyber reaches competitive performance at the frontier on the popular benchmark CyberGym.
5 Flash Cyber. The model will be exclusively available to governments and trusted partners via CodeMender soon as part of a limited-access pilot program. This will give frontline defenders a head start in finding and fixing critical vulnerabilities before they can be exploited, while mitigating against broader misuse. 5 Pro soon. Your information will be used in accordance with Google's privacy policy. You may opt out at any time.



