Google has launched two new speech-generation models, Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS. Flash TTS targets creative direction and character design, while Flash-Lite is positioned for higher-volume, cost-sensitive workloads such as dubbing, content production and voice agents.
The headline capability is prompt-based voice design. Instead of starting with a small set of fixed presets, developers can describe the role, accent, vocal characteristics and delivery they want. Google says the models cover more than 100 languages and dialects and are accompanied by a library of more than 2,000 production-ready voices.
What happened
Gemini 3.8 Flash TTS can generate new voice profiles from natural-language instructions and can replicate a consistent voice from a 30-second reference sample when the user has the right to use that voice. Google says the replication flow includes consent verification tied to the reference speaker.
Both models support line-by-line performance direction, including pacing, emotional tone, accent shifts and conversational cues. Google also highlights long-form generation with limited speaker drift and native two-speaker scene staging for scripted dialogue.

Where it is available
Gemini 3.8 Flash TTS is rolling out in the Gemini API and Google AI Studio for developers and in Gemini Notebook for end users. Flash-Lite is also rolling out through the API and AI Studio, while Google Vids is the consumer-facing destination named for the lighter model. Gemini Enterprise support is planned through a forthcoming API rollout.
The two-model split is deliberate. Flash TTS is intended for use cases where direction and character consistency matter, while Flash-Lite is tuned for throughput and lower-cost production at scale.
How Google is handling voice replication
Voice replication creates obvious impersonation risks, so Google is pairing the feature with consent checks and provenance tooling. The company says generated audio is watermarked with SynthID, and it also points to C2PA credentials as another layer for tracking content origin.

Why it matters
Speech products have often split into two categories: fast text-to-speech from preset voices and separate, more specialized systems for custom voice creation. Gemini 3.8 brings those workflows closer together, combining voice design, performance direction and scalable API access in one family.
That makes voice a more programmable part of an application. If the promised consistency holds up in production, the models could reduce the amount of separate tooling needed for localized games, learning products, podcasts, dubbing pipelines and conversational agents.


