Google introduced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS. The company called these its most expressive audio generation tools to date. Both models began rolling out immediately in the Gemini API and Google AI Studio.
Flash TTS is positioned for deep creative direction. Flash-Lite TTS targets high-volume, cost-efficient applications. The two models work as complementary tools.
Custom Voice Creation Across 100+ Languages
Gemini 3.8 Flash TTS lets users create original voices from scratch. Natural language prompts specify role, accent, and vocal characteristics. Support covers more than 100 languages and dialects.
This release expands Google's voice library from 30 original voices to more than 2,000 production-ready options. Regional varieties include Mexican Spanish, Quebec French, and Scots English.
A voice replication feature can recreate a vocal profile from a 30-second audio sample. Users must have rights to the voice. The system requires a verbal consent recording from the voice owner. All generated audio carries SynthID watermarking and C2PA credentials.
Voice replication through AI Studio is not available in Illinois, Texas, the European Economic Area, the United Kingdom, Switzerland, and India.
Performance Features for Long-Form and Multi-Speaker Audio
Both models support line-by-line performance direction. They handle long-form generation across hours of audio. Native two-speaker scene staging is available for podcasts and dramatic storytelling.
Non-verbal cues such as laughter, sighs, and active-listening interjections can be scripted into dialogue.
Benchmark Results
Gemini 3.8 Flash TTS took the top position on Hume AI's Voice Design Benchmark with a score of 71.4. It also led accent modeling at 60.8. The two models hold the first and second spots on Hume AI's Overall Quality Index.
In blind human preference evaluations on Voice Arena, the models ranked at the top in languages including Japanese, Brazilian Portuguese, Vietnamese, and Hindi.
Availability and Partner Integrations
Developer platforms Agora, LiveKit, Pipecat, and Vercel support the models through the Gemini API. Companies including Figma, HeyGen, Linguana, Wondercraft, 99.co, and Ollang are integrating the TTS models for dubbing, media localization, and voice agent applications.
Enterprise access via Gemini Enterprise is listed as coming soon.

