Model launches

AI-generated text

Google launches Gemini 3.7 Flash with coding and agent improvements and a temporary API price cut

Google released Gemini 3.7 Flash, a workhorse model optimized for coding, agentic workflows and enterprise knowledge work, three weeks after Gemini 3.6 Flash.

Google launches Gemini 3.7 Flash with coding and agent improvements and a temporary API price cut

Google has announced Gemini 3.7 Flash, the latest iteration in its “workhorse” line, optimized for coding, agentic workflows and knowledge work. The update arrives three weeks after Gemini 3.6 Flash; Google attributes the short interval to developer feedback and algorithmic improvements.

Pricing and temporary discount

Google is launching Gemini 3.7 Flash with introductory API pricing through December 31, 2026: $0.75 per million input tokens and $3.75 per million output tokens. Context caching costs $0.075 per million tokens during this period. On January 1, 2027, standard prices will double to $1.50 per million input tokens and $7.50 per million output tokens, with context caching rising to $0.15 per million tokens. The cut is therefore temporary, but gives teams several months to assess whether Google’s claimed reductions in retries and manual oversight translate into lower total operating costs.

What Google says improved

Google describes Gemini 3.7 Flash as its "most intelligent workhorse model yet for coding and agents." The company says the model better adapts when it encounters roadblocks, clarifies user intent when needed and follows instructions more faithfully. It also purportedly applies more effort to multi-step planning and tool calls, aiming for more disciplined execution with fewer retries and less manual supervision.

A Google DeepMind post accompanying the release reports gains in debugging and issue resolution, the ability to generate more functional web layouts and applications with fewer prompts, and improved reasoning and accuracy on real-world business workflows.

Benchmarks — notable coding improvements

Google’s benchmarks show a substantial generational improvement in several software engineering tests:

  • FrontierCode 1.1 Main (production code quality): Gemini 3.7 Flash 43.6% vs. Gemini 3.6 Flash 34.4%. That narrowly exceeds Claude Sonnet 5 at 42.7% and GPT-5.6 Terra at 41.3% in Google’s table.
  • DeepSWE v1.1 (long-horizon software engineering): 3.7 Flash 65.3% vs. 3.6 Flash 49.0%. GPT-5.6 Terra is ahead at 69.6% in Google’s comparisons.
  • Web development (Code Arena Elo): 3.7 Flash 1588, 3.6 Flash 1538, Claude Sonnet 5 1541, GPT-5.6 Terra 1523. Google says the new model produces more functional layouts and feature-complete apps in fewer prompts and more closely follows reference screenshots and design systems.

However, the broader benchmark table is mixed. On Terminal-bench 2.1, Gemini 3.7 Flash scores 85.8% vs. GPT-5.6 Terra’s 87.4%. Claude Sonnet 5 leads some multimodal desktop and OS tasks. These results suggest that 3.7 Flash has become substantially more competitive in coding and agent workloads while occupying a lower price tier, rather than universally displacing higher-priced competitors.

Enterprise workflows and document understanding

Improvements extend beyond software development. On AutomationBench (enterprise workflow automation), Gemini 3.7 Flash scores 30.4%, up from 17.0% for 3.6 Flash; Google lists Claude Sonnet 5 at 10.7% and GPT-5.6 Terra at 23.6% there. On GDP.PDF (complex PDF comprehension), 3.7 Flash reaches 34.0% vs. 22.0% for 3.6 Flash.

Those gains matter for enterprise agents, which must often parse long reports, extract relevant information, pick tools to invoke, update systems and produce reviewable documents. Reliability across that chain can matter more than isolated benchmark performance.

Google is applying the model within Gemini Spark: Google AI Pro and Ultra subscribers can use 3.7 Flash in Spark, the company’s personal AI agent, where Google says the upgrade improves knowledge work and tool use across Google Workspace applications. For enterprises, 3.7 Flash is available via the Gemini Enterprise Agent Platform and the Gemini Enterprise app. Google also ships updated safeguards covering chemical, biological, radiological and nuclear risks and cyber-offense misuse.

Pricing as a competitive lever

The introductory pricing is a strategic bid to embed the model into enterprise workflows. For autonomous agents, a single user request can trigger long sequences of model calls, reasoning tokens and tool interactions, so per-token price interacts with model reliability to determine overall cost. A cheaper token price that requires more retries may not save money; conversely, lower introductory token pricing plus improved first-pass accuracy could materially lower the cost of running high-volume coding or document-processing agents.

For context, Gemini 3.6 Flash’s standard API pricing is $1.50 per million input tokens and $7.50 per million output tokens. Google’s benchmark table lists Claude Sonnet 5 at $2 and $10, respectively, and GPT-5.6 Terra at $2 and $12.

Organization, flagship delays and the broader AI picture

The 3.7 Flash launch also highlights internal timing questions at Google. The company has not provided a release date for Gemini 3.5 Pro in Thursday’s announcement, despite earlier statements that the flagship would arrive months ago; Google’s latest broadly available Pro model remains Gemini 3.1 Pro from February. Reuters reported that Gemini 3.5 Pro missed its target after falling short of internal goals, particularly in coding, even as Google trains Gemini 4.

Leadership changes have occurred alongside these model developments. Demis Hassabis relinquished day-to-day control of DeepMind to become its chair and Alphabet’s chief scientist. Koray Kavukcuoglu now runs DeepMind as a senior vice president reporting directly to CEO Sundar Pichai and oversees Gemini model development, frontier research, the Gemini app and developer teams. Several senior figures, including Jeff Dean and Noam Shazeer, have left to found or join other ventures.

Analyses differ on the implications: some argue Google increasingly prioritizes profitable cloud infrastructure over keeping its own models at the absolute frontier; others point to Google’s large ecosystem advantages (Search, Workspace, Android, Cloud, custom AI chips and consumer reach). Google says the Gemini app has surpassed 950 million monthly users. Current benchmarks place Google behind some leaders overall but still competitive — Artificial Analysis reports scores that place Claude Opus 5 ahead while Google’s Intelligence Index score for Gemini 3.7 Flash improved from 52 to 56.

Availability and developer considerations

Developers can access Gemini 3.7 Flash via the Gemini API in Google AI Studio and Android Studio, and in Google’s Antigravity environment. Enterprises can deploy it through the Gemini Enterprise Agent Platform and Gemini Enterprise. The rapid three-week jump from 3.6 Flash to 3.7 Flash suggests Google can push algorithmic improvements into production without waiting for a new flagship generation, but production teams still need to benchmark new releases against their own repositories, prompts and tool schemas before changing deployments.

Conclusion

Gemini 3.7 Flash shows significant improvements in production coding, web development, document comprehension and workflow automation in Google’s measurements, while its introductory pricing makes it an attractive option for high-volume agent deployments. The model’s long-term advantage will depend on whether these gains reduce the real cost per successfully completed task, particularly after prices return to their standard levels in January 2027.