Google has launched two new speech-generation models, Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS. Flash TTS targets creative direction and character design, while Flash-Lite is positioned for higher-volume, cost-sensitive workloads such as dubbing, content production and voice agents.

The headline capability is prompt-based voice design. Instead of starting with a small set of fixed presets, developers can describe the role, accent, vocal characteristics and delivery they want. Google says the models cover more than 100 languages and dialects and are accompanied by a library of more than 2,000 production-ready voices.

What happened

Gemini 3.8 Flash TTS can generate new voice profiles from natural-language instructions and can replicate a consistent voice from a 30-second reference sample when the user has the right to use that voice. Google says the replication flow includes consent verification tied to the reference speaker.

Both models support line-by-line performance direction, including pacing, emotional tone, accent shifts and conversational cues. Google also highlights long-form generation with limited speaker drift and native two-speaker scene staging for scripted dialogue.

A studio microphone, headphones and an empty music stand inside a professional vocal booth.
Google’s new models add finer control over synthesized performance and custom voice design.

Where it is available

Gemini 3.8 Flash TTS is rolling out in the Gemini API and Google AI Studio for developers and in Gemini Notebook for end users. Flash-Lite is also rolling out through the API and AI Studio, while Google Vids is the consumer-facing destination named for the lighter model. Gemini Enterprise support is planned through a forthcoming API rollout.

The two-model split is deliberate. Flash TTS is intended for use cases where direction and character consistency matter, while Flash-Lite is tuned for throughput and lower-cost production at scale.

How Google is handling voice replication

Voice replication creates obvious impersonation risks, so Google is pairing the feature with consent checks and provenance tooling. The company says generated audio is watermarked with SynthID, and it also points to C2PA credentials as another layer for tracking content origin.

An audio engineer workstation with multitrack audio, a microphone and a separate vocal booth.
Gemini 3.8 TTS is positioned for both directed voice production and scalable speech workflows.

Why it matters

Speech products have often split into two categories: fast text-to-speech from preset voices and separate, more specialized systems for custom voice creation. Gemini 3.8 brings those workflows closer together, combining voice design, performance direction and scalable API access in one family.

That makes voice a more programmable part of an application. If the promised consistency holds up in production, the models could reduce the amount of separate tooling needed for localized games, learning products, podcasts, dubbing pipelines and conversational agents.

Sources

  1. Google — Gemini 3.8 text-to-speech says hello
  2. Google DeepMind — Gemini Audio
  3. Unite.AI — Google Rolls Out Gemini 3.8 Speech Models In API And AI Studio