Key takeaways
- Google is expanding the Gemini lineup with two Flash models and a specialized Cyber version.
- 6 Flash consistently beats in benchmarks.
- Despite gains on multimodal tasks and a one million token context window, Google still trails the best models from competitors in the US and China.
What happened
Google is expanding the Gemini lineup with two Flash models and a specialized Cyber version. 5 Pro, is still missing. 5 Flash Cyber. " Google says pretraining for Gemini 4 is already underway. " That reads like damage control, and it makes clear that Google knows what the market expects but can't deliver yet. 5 Flash. Google says the savings reach 65 percent on specific benchmarks such as DeepSWE.
The company later added native computer use, allowing the model to operate browsers, desktops, and mobile devices on its own. 5 Flash and tuned for cybersecurity work. Google has built it into CodeMender, Google DeepMind's code security agent. Several Flash Cyber subagents work in parallel and combine their results into one report. 6 percent, despite being a much smaller model. 6.
Why it matters
6 Flash consistently beats in benchmarks. 5 Flash. 4 to 83 percent. The GDPval-AA v2 knowledge work benchmark improves from 1,349 to 1,421 points. Computer Use is now a built-in client-side tool in the Gemini API and Gemini Enterprise. Google has also added stronger Frontier Safety safeguards against CBRN misuse and cyberattacks. CBRN refers to chemical, biological, radiological, and nuclear threats.
Despite gains on multimodal tasks and a one million token context window, Google still trails the best models from competitors in the US and China. Logan Kilpatrick, a member of the technical staff, responded to criticism on X, saying the explicit goal was efficiency, usability, and lower cost, and that performance still improved in the process. 5 Flash-Lite is tuned for low latency and high throughput.
According to Artificial Analysis, it produces 350 output tokens per second. 50 per million output tokens. Google says Flash-Lite beats the older 3 Flash on several agentic and coding benchmarks, including SWE-Bench Pro and OSWorld-Verified. 1 score rises from 31 to 54 percent. 5 Flash at its last I/O conference as the centerpiece of its agent strategy.
What to watch
When scanning commits in the V8 JavaScript engine, Flash Cyber turned up 55 confirmed unique findings. 6 found 36, and Google says ten of Flash Cyber's findings didn't show up in any other model's results. In another test, Google's Cloud Vulnerability Research Team used the model to scan public APIs. It found remote code execution flaws within two hours and produced a working exploit that bypassed security protections.
Google says the model is just as useful for offense as it is for defense, so the company is keeping access tight. 5 Flash Cyber through CodeMender as part of a pilot p




