Model launches

Google launches token-efficient Gemini Flash models to lower AI agent costs

Google DeepMind introduced three proprietary Gemini Flash models—Gemini 3.6 Flash, Gemini 3.5 Flash-Lite and the specialized Gemini 3.5 Flash Cyber—aimed at reducing token usage and latency for enterprise AI agents.

Google launches token-efficient Gemini Flash models to lower AI agent costs

Google DeepMind on Thursday released three new proprietary Gemini models that the company describes as among its most token‑efficient yet: Gemini 3.6 Flash, Gemini 3.5 Flash‑Lite and a Cyber‑tuned Gemini 3.5 Flash Cyber. The trio is aimed at making enterprise AI agents faster, smarter and less expensive to run at scale.

Pricing and availability

Gemini 3.6 Flash and Gemini 3.5 Flash‑Lite are available immediately via the Gemini API in Google AI Studio and Android Studio, and are also integrated into the consumer Gemini app and Google Search. Google prices Gemini 3.6 Flash at $1.50 per one million input tokens and $7.50 per one million output tokens through its API. Gemini 3.5 Flash‑Lite is offered at $0.30/$2.50 per million tokens in/out.

Google has not published a specific per‑token price for Gemini 3.5 Flash Cyber; the company said the cybersecurity‑focused model will be “exclusively available to governments and trusted partners via CodeMender soon.”

For comparison, the earlier Gemini 3.5 Flash was priced at $1.50/$9.00 per million tokens, while Gemini 3.1 Flash‑Lite remains listed at $0.25/$1.50 per million tokens and is described by Google as its most cost‑efficient model despite being roughly twice as slow as the new 3.5 Flash‑Lite.

Model limits and knowledge cutoff

According to Google’s official model cards, both Gemini 3.6 Flash and Gemini 3.5 Flash‑Lite support a 1,000,000‑token input context window and a maximum output limit of 64,000 tokens. Both models share a knowledge cutoff date of March 2026.

Token efficiency and benchmark performance

Google and independent benchmarks report substantial token‑efficiency gains. The Artificial Analysis Index indicates Gemini 3.6 Flash reduces output token usage by about 17% relative to Gemini 3.5 Flash. In long‑horizon software engineering benchmarks such as DeepSWE, token savings can reach up to 65%.

Measured performance improvements include:

  • DeepSWE: Gemini 3.6 Flash scores 49% compared with 37% for the previous 3.5 version.
  • MLE‑Bench: 63.9% for Gemini 3.6 Flash versus 49.7% previously.
  • OSWorld‑Verified: 83.0% up from 78.4% (Google also highlights client‑side computer use integration via the Gemini API and Gemini Enterprise).
  • GDPval‑AA v2: score increased from 1,349 to 1,421.

Google states the models reach these results by taking fewer reasoning steps and making fewer tool calls, reducing verbosity and therefore overall token consumption — an efficiency likened to better fuel economy in a vehicle.

Speed and targeted use cases

Gemini 3.5 Flash‑Lite is positioned as the fastest member of the 3.5 family: Artificial Analysis measures it processing roughly 350 output tokens per second, about twice the speed of the prior generation Gemini 3.1 Flash‑Lite. This throughput makes it suitable for agentic search, massive document processing and other high‑volume, low‑latency workloads. Developers can tune the model for minimal latency and low internal reasoning or for higher thinking levels when complex subagent workflows are required.

Gemini 3.6 Flash is aimed at heavier duty tasks — complex code migrations, long‑form knowledge work, multimodal processing and demanding enterprise workflows such as in‑depth document parsing and data analysis. Gemini 3.5 Flash Cyber has been fine‑tuned for vulnerability discovery and patching and is integrated with Google’s CodeMender agent; multiple concurrent Cyber agents can collaborate to produce comprehensive vulnerability reports and the model reaches competitive results on the CyberGym benchmark.

Safety, licensing and access restrictions

Google says it applied enhanced Frontier Safety protections to harden the Flash models against jailbreaks and to reduce misuse in Chemical, Biological, Radiological and Nuclear (CBRN) domains and cyber‑offensive applications. The engineering approach balances minimizing harmful uses with avoiding unnecessary refusals for legitimate tasks.

All three Flash models are proprietary and available only through Google’s commercial APIs; Google does not publish model weights, training data or source code. This commercial, closed licensing means developers and enterprises effectively rent inference access under Google’s terms, metering and pricing. Full air‑gapped local deployment without special enterprise arrangements is not supported for typical customers.

The Gemini 3.5 Flash Cyber model has further access restrictions: it will be distributed via a limited pilot program to governments and trusted partners to mitigate dual‑use risks.

Questions about a flagship model and future work

Some developers have asked about the larger Gemini 3.5 Pro model Google previously indicated would arrive this summer. Gemini 3.1 Pro debuted in February 2026; Logan Kilpatrick of Google responded on X that Gemini 3.5 Pro is testing with partners and will be made broadly available once ready. Google also confirmed that pre‑training for Gemini 4 has begun.

Bottom line

The Gemini Flash series emphasizes token efficiency and latency reductions to lower operational costs for agentic, coding and cybersecurity workloads. While the models deliver measurable gains on third‑party benchmarks and offer lower API prices, their proprietary API‑only licensing and restricted access for the cyber model underline Google’s focus on control and safety rather than open‑source distribution. Developers and enterprises will need to weigh the cost, speed and access tradeoffs while awaiting larger flagship releases.