Key takeaways
- Building on the momentum of 3.7 Flash from three weeks ago and marking our third Flash release in only six weeks, today we’re introducing…
- 7 Flash, often approaching the performance of higher-cost frontier models.
- 9% on HLE-Verified, demonstrating its ability to handle multi-step reasoning across STEM, humanities, and professional fields.
What happened
7. 8 introduces 2 variants: While tailored for different deployment environments, both of today's releases are powered by the same foundational intelligence, and further accelerated by long-running agentic loops designed to recursively evaluate and refine the underlying models. The significant coding and reasoning gains across this shared core were driven by a number of innovations, including rigorous training in the highly demanding domain of cybersecurity.
This is why we have invested in vulnerability fixing from the start, and prioritized it over offensive capabilities like exploitation. CWE-Bench, run by Collinear, is a challenging external benchmark for patching capabilities. 8%, yet offered at a significantly lower cost. 8 Flash Cyber to secure code across Google.
Why it matters
7 Flash, often approaching the performance of higher-cost frontier models. 8 Flash outperforms most larger frontier models in autonomously solving complex engineering problems end to end, only at a fraction of the cost. 8 Flash exhibits the dependability required for critical enterprise autonomy, across specialized knowledge domains. 7 Flash and other frontier models in benchmarks like Vals Finance Agent V2 and Harvey's Legal Agent Benchmark.
9% on HLE-Verified, demonstrating its ability to handle multi-step reasoning across STEM, humanities, and professional fields. 8 Flash works harder. On complex tasks, it exhibits greater diligence — executing extra reasoning steps, and calling tools iteratively. At times, the model might use more tokens to maximize performance, especially at higher effort levels. 7 Flash, which remains fully supported for efficiency-first workloads.
8 Flash Cyber, available to a set of trusted defenders via the Fairwind Program, provides a decisive advantage in today’s complex cybersecurity landscape, with the Flash speed and cost that enables quick iteration. 8 Flash Cyber demonstrates frontier-level performance in autonomous vulnerability discovery. 5 Flash Cyber as well as significantly larger frontier models.
8 Flash Cyber against a comprehensive internal benchmark in which the model has to discover a wide range of vulnerabilities across complex codebases spanning 20 programming languages. Here, the model showcases an impressive leap over our previous models and reaches a success rate exceeding 70%. 8 Flash Cyber, we focused specifically on equipping defenders with expert capabilities that give them an advantage over attackers.
What to watch
8 Flash ships with safeguards against misuse in the domains of Chemical, Biological, Radiological, and Nuclear (CBRN) and cyber offense, while enabling beneficial use cases, as per our Frontier Safety Framework. 8 Flash Cyber ships with a more permissive set of mitigations for cybersecurity, and as such, is only available to trusted defenders who require a more comprehensive set of cyber capabilities.
8 models have also made a significant leap in prompt injection robustness as measured by Gray Swan, protecting Gemini model users from prompt-injection related malicious attacks.




