Introducing Gemini 3.8 Flash and 3.8 Flash Cyber
Google DeepMind has launched Gemini 3.8 Flash and 3.8 Flash Cyber, enhancing AI capabilities for agentic workflows, software engineering, and cybersecurity with improved reasoning and coding at the same low cost.
Intelligence analysis by Gemini 2.5 Flash

The latest Gemini models, 3.8 Flash and 3.8 Flash Cyber, represent Google DeepMind's third Flash release in six weeks, offering significant advancements in complex reasoning, autonomous software engineering, and frontier-level cybersecurity performance. These models are designed for efficiency and intelligence, catering to both general enterprise applications and specialized defensive…
Imagine you have a super-smart robot helper that's gotten even better at solving tricky puzzles and building things. Google just made two new versions: one, called Flash, is super good at helping engineers write computer code and solve complex problems, like building a whole game from a simple idea. The other, called Flash Cyber, is like a super detective for computers, finding hidden weaknesses in software and fixing them really fast, helping keep our digital world safe from bad guys.
Analysis
Google DeepMind's rapid iteration in its Gemini Flash series culminates in the release of Gemini 3.8 Flash and its specialized variant, 3.8 Flash Cyber. This launch underscores a strategic push towards more intelligent, efficient, and domain-specific AI models. The core advancements stem from rigorous training, including extensive exposure to the demanding field of cybersecurity, which has yielded significant gains in both reasoning and coding capabilities across the shared foundational intelligence of both models. The company emphasizes that these models are designed to work harder on complex tasks, executing more reasoning steps and iteratively calling tools to maximize performance, even if it occasionally means using more tokens.
Gemini 3.8 Flash
Gemini 3.8 Flash is positioned as an intelligent workhorse model, demonstrating substantial improvements over its predecessor, 3.7 Flash. It excels in long-horizon software engineering, often matching or exceeding the performance of more expensive frontier models on benchmarks like DeepSWE v1.1. This model is also engineered for dependability in critical enterprise autonomy, showing superior performance in specialized knowledge domains such as finance and legal analysis, as evidenced by benchmarks like Vals Finance Agent V2 and Harvey's Legal Agent Benchmark. Its ability to handle multi-step reasoning across diverse fields, including STEM and humanities, is highlighted by a 54.9% score on HLE-Verified.
The enhanced performance of 3.8 Flash is attributed to its design philosophy of 'working harder,' which involves executing extra reasoning steps and iterative tool calls. While this approach maximizes performance, developers have the flexibility to adjust 'effort levels' to manage token overhead for applications where compute efficiency is paramount. Google DeepMind showcases the model's versatility through examples like building a 3D wizard game, a functional DOS version of Google Maps, and an interactive 3D visualizer for hardware anatomy, all from simple prompts within Google Antigravity or AI Studio.
CyberGym
Gemini 3.8 Flash Cyber distinguishes itself with frontier-level performance in cybersecurity, specifically in autonomous vulnerability discovery. On CyberGym, a standard industry benchmark for finding vulnerabilities, the model surpasses both its predecessor, 3.5 Flash Cyber, and significantly larger frontier models. This achievement is particularly noteworthy as CyberGym primarily focuses on C/C++ codebases, a critical area for software security.
Beyond C/C++, Google DeepMind also evaluated 3.8 Flash Cyber against an internal benchmark encompassing a wide range of vulnerabilities across complex codebases in 20 programming languages. In this more comprehensive real-world scenario, the model demonstrated an impressive leap, achieving a success rate exceeding 70%. This broad language support and high success rate underscore its potential to provide a decisive advantage to defenders in today's complex and multi-faceted cybersecurity landscape.
Fairwind Program
Gemini 3.8 Flash Cyber is made available to trusted defenders through a new initiative called the Fairwind Program. This program is designed to equip cybersecurity professionals with expert capabilities, focusing on defensive actions like vulnerability fixing rather than offensive exploitation. The emphasis on automated patching is a key differentiator, with the model demonstrating strong performance on challenging external benchmarks like CWE-Bench, run by Collinear.
On CWE-Bench, Gemini 3.8 Flash Cyber achieves a pass@1 score of 47.2%, closely matching a leading frontier model at 47.8%, but at a significantly lower cost. This cost-effectiveness, combined with its speed, enables quick iteration in defensive strategies. Google itself is already leveraging 3.8 Flash Cyber to secure its own code, with the Chrome Security team reporting that the model produced 2.6 times more correct patches for Chrome vulnerabilities compared to the best commercial models, which are often much larger.
Key points
- Google DeepMind released Gemini 3.8 Flash and 3.8 Flash Cyber, marking the third Flash release in six weeks.
- Gemini 3.8 Flash offers significant improvements in software engineering, agentic tasks, and multi-step reasoning at the same low cost as 3.7 Flash.
- Gemini 3.8 Flash Cyber provides frontier-level performance in autonomous vulnerability discovery and automated patching for cybersecurity.
- Both models are powered by the same foundational intelligence, enhanced by rigorous training, including in cybersecurity.
- Gemini 3.8 Flash Cyber is available to trusted defenders through the new Fairwind Program and is already used to secure Google's own code.
The introduction of Gemini 3.8 Flash and 3.8 Flash Cyber promises to significantly accelerate software development and enhance cybersecurity defenses. Their improved reasoning and coding capabilities, coupled with cost-efficiency, could lead to more robust and secure digital systems across various industries, empowering developers and defenders alike.
While highly capable, the models' tendency to use more tokens for maximum performance could lead to increased operational costs for some applications, potentially limiting their adoption in highly cost-sensitive environments. Additionally, the reliance on AI for critical security functions introduces new layers of complexity and potential for unforeseen vulnerabilities if not meticulously managed.


