Today, Google announced two new text‑to‑speech (TTS) models in the Gemini family: Gemini 3.8 Flash and Gemini 3.8 Flash‑Lite. The company frames the release as a move from static preset voices toward a dynamic voice design workspace that lets creators, developers and enterprises build richer, more expressive audio experiences.
Model focus and intended uses
-
Gemini 3.8 Flash TTS: designed for deep creative direction and character design. Creators can generate entirely new voices from natural language prompts and control performance at a granular, line‑by‑line level (acting cues, pacing, dialect shifts, backchanneling). Use cases include games, immersive audiobooks, podcasts and interactive media.
-
Gemini 3.8 Flash‑Lite TTS: optimized for high‑volume, cost‑efficient scaling. Targeted at large‑scale dubbing, audio content production and expressive voice agents, it provides fine‑grained control over tone, pacing and expressive nuance while prioritizing throughput and cost.
Generative voice design and voice library
With Gemini 3.8 Flash, users can scale beyond 30 original voices to an effectively infinite library: role, accent and voice characteristics can be customized via natural language prompts across more than 100 languages and dialects. The offering includes over 2,000 production‑ready voices and covers regional variants such as Mexican Spanish, Quebec French and Scots English.
Voice replication is supported from as little as a 30‑second audio sample for voices users have rights to, with built‑in consent verification, SynthID watermarking and C2PA credentials intended to protect developers and talent. A voice remixing feature is planned that will let users adjust timbre, pitch, pace and accent on library voices using prompts.
Line‑by‑line performance control
Both models provide precise control over line delivery. Creators can include stage directions in scripts or rely on Gemini’s natural script cues to steer delivery — from calm customer service agents to whispery suspense scenes. The models support long‑form generation with stable timbre and pacing over hours of audio, native two‑speaker scene staging from a single script, and scripted non‑verbal cues and backchanneling (e.g., <laughs>, <sigh>, |mhm|) to add conversational texture.
Performance metrics and language coverage
Google cites Hume AI benchmark results: Gemini 3.8 Flash TTS scored 71.4 on the Voice Design Benchmark and 60.8 in accent modeling. On Hume AI’s Overall Quality Index, Gemini 3.8 Flash and Flash‑Lite rank first and second, respectively. In blind human preference tests on Voice Arena, the new models placed highly in several global languages including Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic, Mexican Spanish and Hindi. The models support over 100 languages.
Trust, consent and transparency
Google says it designed voice creation and replication with safeguards to protect voice talent and identity. Voice replication requires a verbal consent recording from the voice owner that matches the reference speaker. All audio produced by Gemini Audio models is watermarked with SynthID, an imperceptible audio watermark intended to help detect AI‑generated speech. Further safety and responsibility details are provided in the model card.
Availability and integrations
Developers can try the new speech generation features today in Google AI Studio. The Gemini API enables developer platforms such as Agora, LiveKit, Pipecat and Vercel to build and deploy speech generation experiences. Google lists partners integrating the TTS models, including Figma, HeyGen, Linguana, Wondercraft, 99.co and Ollang, for use cases such as global dubbing, localized regional accents and large‑scale conversational agents.
Rollout details:
- Gemini 3.8 Flash TTS: available today for developers via the Gemini API and Google AI Studio; enterprise API access coming soon via Gemini Enterprise; will be available to everyone in Gemini Notebook.
- Gemini 3.8 Flash‑Lite TTS: available today for developers via the Gemini API and Google AI Studio; enterprise API access coming soon via Gemini Enterprise; will be available to everyone in Google Vids.
A notable restriction: voice replication via AI Studio is not available in Illinois, Texas, the EEA, the UK, Switzerland and India.
Summary
Gemini 3.8 Flash and Flash‑Lite aim to provide both deep creative voice design (Flash) and cost‑effective, high‑volume speech generation (Flash‑Lite). Google highlights broad language support, advanced customization features and built‑in safety mechanisms, with developer access available today through the Gemini API and Google AI Studio.



