Google has announced three additions to its Gemini Flash family aimed at scaling production agent workflows: Gemini 3.6 Flash, Gemini 3.5 Flash‑Lite, and a cyber‑focused Gemini 3.5 Flash Cyber that will be offered through CodeMender.
Why this matters
The company says developers and enterprise customers running agentic systems need better token efficiency, lower latency, and more reliable performance. The Flash line is positioned to balance efficiency and quality so agents become cheaper and faster to operate.
Gemini 3.6 Flash: more efficient and higher quality than 3.5 Flash
Gemini 3.6 Flash builds on feedback from 3.5 Flash and is aimed at improvements in coding, knowledge work and multimodal tasks. According to the Artificial Analysis Index cited by Google, 3.6 Flash uses 17% fewer output tokens than 3.5 Flash; in some benchmarks such as Datacurve’s DeepSWE, token reductions of up to 65% were observed. The model also requires fewer reasoning steps and tool calls for multi‑step workflows.
Pricing (announced):
- Input tokens: $1.50 / 1M
- Output tokens: $7.50 / 1M Google says this translates into a lower overall cost per agent task compared with 3.5 Flash.
Benchmark and capability highlights (as reported):
- DeepSWE (Datacurve): fewer unwanted code edits and reduced execution loops (49% vs. 37%).
- MLE Bench (ML research): 63.9% vs. 49.7%.
- OSWorld‑Verified (computer use): 83.0% vs. 78.4%.
- GDPval‑AA v2 (knowledge work): 1421 vs. 1349.
Example customer use cases provided by Google:
- Parsing and analysing financial data and transcripts with Managed Agents on AIS more efficiently than 3.5 Flash.
- Performing code migrations with multi‑agent orchestration on AGY with lower latency and higher quality.
- Developing a photographic texture extractor for 3D workflows using the Gemini App canvas.
- Building interactive theme studios using AGY and the tldraw offline editor thanks to stronger visual understanding.
Safety measures: 3.6 Flash ships with enhanced Frontier Safety safeguards for CBRN (chemical, biological, radiological, nuclear) and cyber‑offense misuse. Google says these measures make the model substantially more resistant to jailbreaks while reducing unnecessary refusals for beneficial uses. More details are available in the 3.6 Flash model card.
Gemini 3.5 Flash‑Lite: designed to scale agentic workflows
Gemini 3.5 Flash‑Lite is positioned as the fastest model in the 3.5 series, optimized for low‑latency and high‑throughput developer workloads such as agentic search and document processing. Artificial Analysis reports it runs at 350 output tokens per second.
Pricing (announced):
- Input tokens: $0.30 / 1M
- Output tokens: $2.50 / 1M
Performance improvements over prior Flash‑Lite generations and other models:
- Terminal‑Bench 2.1 (coding/agent tasks): 54% vs. 31% (vs. 3.1 Flash‑Lite).
- GDM‑MRCR v2 (long context): 72.2% vs. 60.1%.
- GDPval‑AA v2 (real‑world task execution): 1140 vs. 642.
Google also reports that 3.5 Flash‑Lite outperforms 3 Flash on several agentic and coding evaluations (for example, SWE‑Bench Pro and OSWorld‑Verified), making it a faster and more capable option for many 2.5 and 3 Flash workloads. The model includes a built‑in computer use tool to support agentic tasks across surfaces.
Representative applications mentioned:
- Extracting product features from massive e‑commerce datasets and synthesizing results.
- Generating 25 unique web design concepts instantly when working with 3.6 Flash as a master agent.
- Scaling receipt translation and summarization using its multimodal understanding.
- Rapidly generating and iterating game design options.
Early customers highlight the model’s combination of speed, intelligence and cost efficiency for scaling agentic workflows and data processing.
Gemini 3.5 Flash Cyber in CodeMender: detecting and fixing vulnerabilities
Gemini 3.5 Flash Cyber is a security‑specialized variant built on 3.5 Flash and fine‑tuned for finding and fixing cybersecurity vulnerabilities at a lower token cost than larger models. Within CodeMender, multiple 3.5 Flash Cyber agents collaborate to produce combined reports; Google says the approach reaches competitive, frontier‑level performance on the CyberGym benchmark.
Because of dual‑use concerns, Google is adopting a restricted deployment strategy: 3.5 Flash Cyber will be available exclusively to governments and trusted partners through CodeMender in a limited‑access pilot. The intent is to give frontline defenders tools to find and patch critical vulnerabilities while limiting broader misuse.
Availability and future work
Google states that Gemini 3.6 Flash and 3.5 Flash‑Lite are available starting today:
- For developers via the Gemini API in Google AI Studio and Android Studio; 3.6 Flash is also available in Google Antigravity. A Developer Guide is provided.
- For enterprises on the Gemini Enterprise Agent Platform; 3.6 Flash is available in the Gemini Enterprise app.
- For general users via the Gemini app; 3.5 Flash‑Lite is also rolling out in Google Search.
The company also said Gemini 3.5 Pro is currently being tested with partners and will be made generally available when ready. In parallel, Google has started its largest pre‑training run yet for Gemini 4 and is continuing work on the next generation of models.
Google invited developers and partners to provide feedback on 3.6 Flash and 3.5 Flash‑Lite to inform future Gemini releases.



