Google Gemini 3.8 Flash and Flash-Lite TTS introduces expressive AI voice generation with custom voices, natural speech controls, multilingual support and long-form audio capabilities...
- Google Gemini 3.8 TTS
- Flash and Flash-Lite Models
- Custom Voice Creation
- Voice Replication
- Expressive Voice Controls
- Multilingual Audio Support
- Long-Form Audio Generation
- Two-Speaker Conversations
- Availability and Safety
Google has introduced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, two new text-to-speech models designed to generate more expressive and customizable AI voices. Google describes them as its most expressive audio-generation models yet.
Gemini 3.8 Flash TTS
This is the more advanced model focused on creative voice generation and detailed performance control.
It can create completely new voices from natural-language descriptions. Developers can specify characteristics such as:
- Character and personality
- Accent
- Vocal style
- Emotion
- Pacing
- Acting direction
- Dialect
Google says the model supports more than 100 languages and dialects.
Gemini 3.8 Flash-Lite TTS
The Flash-Lite version is designed for high-volume and cost-efficient speech generation.
Google positions it for applications such as large-scale dubbing, audio content creation and expressive voice agents. It still provides controls for tone, pacing and other vocal characteristics.
Custom Voice Creation
One of the major features is the ability to design a voice from scratch using a natural-language prompt.
Google says Gemini 3.8 Flash TTS provides access to 2,000+ production-ready voices, while developers can also create custom vocal identities.
Voice Replication
The system can replicate a voice using approximately a 30-second audio sample, but Google has added consent and safety requirements.
The person whose voice is being replicated must provide a matching verbal consent recording. Google also says generated audio receives SynthID watermarking and C2PA credentials.
Importantly, Google's post says voice replication through AI Studio is currently unavailable in India, along with Illinois, Texas, the EEA, UK and Switzerland.
Line-by-Line Voice Direction
Gemini 3.8 Flash TTS allows creators to direct individual lines using natural-language instructions.
For example, a script can specify different delivery styles, including calm speech, suspenseful delivery, whispers, emotional reactions and conversational responses.
It also supports vocal effects and conversational sounds such as laughter, sighs, gasps and backchannel responses like “mhm” or “yeah.”
Long-Form Audio
Google says the models are designed to maintain voice quality, natural pacing and character consistency during hours of continuous audio generation.
That makes the technology particularly relevant to audiobooks, podcasts and long-form narrated content.
Two-Speaker Conversations
The models support native two-speaker scene staging. A single script can contain multi-turn dialogue while keeping the two voices distinct.
This can be useful for podcasts, dramatic storytelling, games and conversational AI applications.
Performance and Language Support
Google reports that Gemini 3.8 Flash TTS achieved the #1 position on Hume AI's Voice Design Benchmark, with a score of 71.4, and a 60.8 score for accent modeling.
Google also says both models performed strongly in Voice Arena evaluations across languages including Hindi, Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic and Mexican Spanish.
These are Google-reported benchmark results, so they should be understood as the company's reported evaluations rather than an independent overall ranking.
Where You Can Use It?
As of the September 23 announcement:
Gemini 3.8 Flash TTS
- Gemini API
- Google AI Studio
- Gemini Notebook
- Gemini Enterprise: coming soon
Gemini 3.8 Flash-Lite TTS
- Gemini API
- Google AI Studio
- Google Vids
- Gemini Enterprise: coming soon
Safety and Transparency
Google says all audio generated by its Gemini Audio models is watermarked with SynthID. The company says the watermark is designed to make AI-generated audio detectable.
For voice replication specifically, consent verification is required before a voice can be created.
Also Read: AI News: OpenAI New Models: GPT-6 Astra and GPT-5.6 Update
What Makes This Release Important?
The key change is that Google's TTS technology is moving beyond simply converting written text into a predefined voice. Gemini 3.8 Flash TTS is designed more like a voice-performance studio, allowing creators to design voices and control how individual lines are acted.
For developers, the release is particularly relevant to AI voice agents, games, audiobooks, podcasts, dubbing, interactive storytelling and multilingual applications.
Also Read: Kimi AI vs ChatGPT 2026: Which AI Assistant Is Better?

