Model launches

Google launches Gemini 3.6 Flash and 3.5 Flash‑Lite; limited 3.5 Flash Cyber pilot for CodeMender

Google introduced two new models in its Flash family — Gemini 3.6 Flash, focused on higher efficiency and quality, and Gemini 3.5 Flash‑Lite, aimed at low latency and high throughput for agentic workflows.

Google launches Gemini 3.6 Flash and 3.5 Flash‑Lite; limited 3.5 Flash Cyber pilot for CodeMender

Google announced new additions to its Gemini Flash family aimed at improving token efficiency, lowering latency, and delivering more reliable performance for production AI agents. The releases include Gemini 3.6 Flash and Gemini 3.5 Flash‑Lite, plus a cyber‑focused Gemini 3.5 Flash Cyber variant that will be offered in a limited pilot via CodeMender.

Gemini 3.6 Flash: improved efficiency and quality

Gemini 3.6 Flash is presented as a workhorse model that enhances coding, knowledge work, and multimodal tasks while improving token efficiency relative to Gemini 3.5 Flash. According to the Artificial Analysis Index, 3.6 Flash uses 17% fewer output tokens than 3.5 Flash on average; in some benchmarks such as DeepSWE measured by Datacurve, Google reports up to 65% reduction in output tokens. The company says these efficiency gains come with a lower cost per output token.

Pricing and cost implications: Google lists 3.6 Flash at $1.50 per 1M input tokens and $7.50 per 1M output tokens, which the company says reduces the overall cost per agentic task compared to 3.5 Flash.

Performance and concrete examples:

  • Higher precision with fewer unwanted code edits and fewer execution loops (DeepSWE figures noted as 49% vs. 37%).
  • Improved results in ML research benchmarks (MLE Bench: 63.9% vs. 49.7%).
  • Better computer‑use capabilities in OSWorld‑Verified (83.0% vs. 78.4%).
  • Stronger performance on knowledge work benchmarks (GDPval‑AA v2 score: 1421 vs. 1349).

Google customers such as Hebbia and Harvey highlighted 3.6 Flash’s strengths on multimodal tasks like document parsing, chart and data analysis, and report drafting. Example applications cited by Google include analyzing financial data and transcripts with Managed Agents, executing code migrations through multi‑agent orchestration, building photographic texture extractors for 3D workflows, and creating interactive theme studios using offline editors.

Safety: 3.6 Flash ships with enhanced Frontier Safety protections addressing Chemical, Biological, Radiological, and Nuclear (CBRN) risks and cyber‑offense misuse. Google says these safeguards make the model substantially more resistant to jailbreak attempts while training it to minimize refusals for beneficial uses. Further details are available in the 3.6 Flash model card.

Gemini 3.5 Flash‑Lite: designed to scale agentic workflows

Gemini 3.5 Flash‑Lite targets low‑latency and high‑throughput developer workloads such as agentic search and document processing. Artificial Analysis reports 3.5 Flash‑Lite runs at 350 output tokens per second, making it the fastest model in the 3.5 class.

Pricing: $0.30 per 1M input tokens and $2.50 per 1M output tokens. Google states that 3.5 Flash‑Lite offers significantly better quality than 3.1 Flash‑Lite while providing a strong price‑to‑performance ratio for production traffic.

Capabilities and benchmarks:

  • Executes high‑volume tasks with lower latency than 3.5 Flash.
  • Scales agentic systems effectively; substantially outperforms 3.1 Flash‑Lite across thinking levels.
  • Shows improvements in coding and agentic tasks (Terminal‑Bench 2.1: 54% vs. 31%), long‑context handling (GDM‑MRCR v2: 72.2% vs. 60.1%), and real‑world task execution (GDPval‑AA v2: 1140 vs. 642).
  • In many agentic and coding evaluations, 3.5 Flash‑Lite even surpasses some 3 Flash results (SWE‑Bench Pro: 54.2% vs. 49.6%; OSWorld‑Verified: 74.0% vs. 65.1%).

Google provided application examples: extracting product features from large e‑commerce datasets, generating web design concepts when paired with 3.6 Flash as a master agent, scaling receipt translation and summarization with multimodal understanding, and rapidly prototyping simple games.

Early customers emphasize the combination of speed, capability, and cost efficiency that enables scaling agentic workflows and data processing.

More details are available in the 3.5 Flash‑Lite model card.

Gemini 3.5 Flash Cyber in CodeMender: targeted vulnerability detection and patching

Google also introduced Gemini 3.5 Flash Cyber, a specialized, efficiency‑focused model built on 3.5 Flash and fine‑tuned for finding and fixing cybersecurity vulnerabilities at lower token cost than larger models. Within CodeMender — where multiple 3.5 Flash Cyber agents collaborate to produce a combined report — the model reached competitive frontier performance on the CyberGym benchmark.

Because the technology has dual‑use risks, Google is deploying 3.5 Flash Cyber intentionally and restrictively: availability will be limited to governments and trusted partners through a CodeMender pilot program. The stated aim is to give frontline defenders an advantage in finding and fixing critical vulnerabilities before they can be exploited, while reducing the risk of broader misuse.

Availability and future work

Google says Gemini 3.6 Flash and 3.5 Flash‑Lite are available starting today on multiple surfaces:

  • For developers via the Gemini API in Google AI Studio and Android Studio. 3.6 Flash is also available in Google Antigravity; developers can consult the Developer Guide to get started.
  • For enterprises via the Gemini Enterprise Agent Platform; 3.6 Flash is additionally available in the Gemini Enterprise app.
  • For general users via the Gemini app; 3.5 Flash‑Lite is also rolling out in Google Search.

Google also noted that Gemini 3.5 Pro is in partner testing and will be made broadly available when ready. The company has begun its largest pretraining run yet for Gemini 4 and continues to seek feedback from developers and customers to inform future model releases.