Google has introduced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, two text-to-speech models built to give users more control over AI-generated audio. The models can create voices from written descriptions, follow detailed delivery directions, generate two-speaker conversations, and maintain consistent voices during longer recordings.

The models can support game characters, podcasts, audiobooks, and video narration. For narration and voiceover work, users can describe a preferred voice and direct how individual lines should sound.

Create Voices From Descriptions

Gemini 3.8 Flash TTS can generate a voice based on a natural-language description. Users can specify a character, accent, and other vocal qualities. They can also select from more than 2,000 existing voices.

The models support more than 100 languages and dialects.

Custom voices can be saved and reused in different projects. This is intended to help a narrator or character keep a consistent sound rather than shifting over time.

Google also plans to introduce voice remixing later. That feature will allow existing voices to be adjusted for:

  • Pitch
  • Pace
  • Timbre
  • Accent

Add Delivery Directions Within a Script

Gemini’s speech models allow instructions to be placed throughout a script, not just set once for an entire recording. A line can be directed to sound whispered, slower, or more emotional.

Scripts can also include vocal and conversational cues, such as:

  • Laughter
  • Sighs
  • Gasps
  • “Mhm”
  • “Yeah”

This gives users a way to guide the delivery of individual lines while creating narration, dialogue, or character-based audio.

Generate Two-Speaker Conversations

Gemini 3.8 Flash TTS can produce a two-speaker conversation from a single script. The system can keep voices and pacing consistent across longer recordings.

That capability can be used where a script requires more than one voice, while still keeping the production inside one text-to-speech workflow.

Voice Replication Requirements and Safeguards

Gemini can recreate a consistent vocal profile from a 30-second recording. However, users cannot submit just any voice sample.

Google states that the recording must belong to the user or be a voice the user has rights to use. Before a replica can be generated, the voice owner must also provide a separate verbal consent recording that matches the reference speaker.

Generated Gemini Audio clips include an imperceptible SynthID watermark. Google also says voice replication includes C2PA credentials to help identify AI-generated material.

Gemini 3.8 Flash TTS Availability

Gemini 3.8 Flash TTS is rolling out through:

  • Gemini Notebook
  • The Gemini API
  • Google AI Studio

Gemini 3.8 Flash-Lite TTS is the lower-cost model. It is coming to Google Vids and is designed for larger-scale work, including dubbing, narration, and voice agents.