Model launches

Google launches token‑efficient Gemini Flash models for agentic and cybersecurity use

Google DeepMind introduced three proprietary Gemini Flash models—Gemini 3.6 Flash, Gemini 3.5 Flash‑Lite, and the specialized Gemini 3.5 Flash Cyber—designed to reduce token usage and latency for agentic workloads.

Google launches token‑efficient Gemini Flash models for agentic and cybersecurity use

Google DeepMind today announced three new proprietary Gemini models that the company describes as among its most token‑efficient to date: Gemini 3.6 Flash, Gemini 3.5 Flash‑Lite, and the specialized Gemini 3.5 Flash Cyber. The lineup aims to improve speed, capability and cost‑efficiency for agentic workloads.

Pricing and availability

  • Gemini 3.6 Flash API price: $1.50 per 1 million input tokens and $7.50 per 1 million output tokens.
  • Gemini 3.5 Flash‑Lite API price: $0.30 per 1 million input tokens and $2.50 per 1 million output tokens.
  • No specific price has been published yet for Gemini 3.5 Flash Cyber; Google says it is tuned to a lower per‑token price than larger models.

Gemini 3.6 Flash and Gemini 3.5 Flash‑Lite are available immediately through the Gemini API in Google AI Studio and Android Studio, and they are integrated into the consumer Gemini app and Google Search. According to Google, Gemini 3.5 Flash Cyber will be made available "exclusively available to governments and trusted partners via CodeMender soon" — CodeMender being Google’s proprietary code bug‑fixing agent released last year.

Context window and token efficiency

Both Gemini 3.6 Flash and Gemini 3.5 Flash‑Lite feature a 1‑million‑token input context window and a maximum output limit of 64,000 tokens. Both models share a knowledge cutoff date of March 2026.

Independent benchmarking group Artificial Analysis reports that Gemini 3.6 Flash reduces output token usage by 17% compared with Gemini 3.5 Flash. On longer‑horizon software‑engineering benchmarks such as DeepSWE, token savings reach up to 65% in some measures. Google states these reductions correspond to fewer reasoning steps and tool calls needed to complete the same multi‑step workflows.

Benchmark improvements

Gemini 3.6 Flash shows concrete gains across several benchmarks: DeepSWE improved from 37% to 49%, MLE‑Bench rose from 49.7% to 63.9%, and OSWorld‑Verified increased from 78.4% to 83.0%. On a knowledge‑work benchmark (GDPval‑AA v2) the score moved from 1349 to 1421. Google attributes these improvements to a combination of reduced verbosity and streamlined internal reasoning.

Roles and capabilities of each model

  • Gemini 3.6 Flash: positioned as the heavy‑duty model for complex coding, multimodal knowledge work and multi‑agent orchestration. Use cases include complex document parsing, advanced chart and data analysis, long‑form report drafting, and large‑scale code migrations with lower latency and higher quality than earlier Gemini versions.

  • Gemini 3.5 Flash‑Lite: targeted at workloads where throughput and minimal latency are critical. Artificial Analysis measures roughly 350 output tokens per second for this model, about twice the speed of the prior Gemini 3.1 Flash‑Lite. It is optimized for agentic search, massive document processing, feature extraction from large datasets and high‑volume tasks where developers can tune thinking levels for latency or complexity.

  • Gemini 3.5 Flash Cyber: a cybersecurity‑focused model fine‑tuned to find and patch vulnerabilities. It integrates with Google’s CodeMender and in practice multiple Cyber agents can collaborate to produce a single comprehensive vulnerability report. The model achieves competitive frontier performance on the CyberGym benchmark.

Pricing context among frontier models

The announced prices place Google’s Flash models in the middle‑to‑lower range compared with other major models. For example, Gemini 3.5 Flash‑Lite’s combined input+output cost is $2.80 per 1M tokens, whereas Gemini 3.6 Flash totals $9.00 per 1M tokens. Google notes that lower internal token usage can further reduce customers’ effective costs since fewer tokens are consumed per task. The company also highlights that Gemini 3.1 Flash‑Lite remains its most cost‑efficient offering at $0.25/$1.50 (input/output), although it is roughly twice as slow as the new 3.5 Flash‑Lite.

Safety and guardrails

Google says it deploys enhanced Frontier Safety safeguards on these models, including strengthened jailbreak protections and mitigations for Chemical, Biological, Radiological and Nuclear (CBRN) risks and cyber‑offense misuse. Engineering choices aim to reduce unnecessary refusals for legitimate uses while maintaining robust safety protections.

Commercial licensing and access restrictions

All three new Gemini models are proprietary and available through Google’s commercial API. The underlying model weights, training data and source code are not publicly released, so developers access the models via metered API calls rather than open‑source licenses like MIT or Apache. This approach limits the ability to self‑host or fully air‑gap the models without specialized enterprise agreements with Google Cloud.

Gemini 3.5 Flash Cyber is even more restricted: Google will initially provide access only to governments and trusted partners through a limited pilot, reflecting concerns around dual‑use cyber capabilities and responsible deployment.

Missing flagship and future roadmap

Developers noted the absence of a broader release of the larger Gemini 3.5 Pro model Google had previously signaled would be released this summer. Gemini 3.1 Pro debuted in February 2026, and competitors have since shipped newer flagship generations. Google technical staffer Logan Kilpatrick replied on X: "Gemini 3.5 Pro is currently testing with partners and we plan to make it broadly available as soon as it’s ready." The company also confirms that pre‑training for Gemini 4 has already started.

Conclusion

The Gemini Flash series signals Google’s bet on more efficient, agentic AI: faster, less verbose models that reduce token consumption and cost for sustained autonomous workflows. The improvements in token efficiency and benchmark performance make these models attractive for enterprises with large‑scale agentic, coding and cybersecurity needs, but their closed, commercial licensing bounds deployment flexibility and keeps advanced capabilities under Google’s access controls.