The search giant has released two new Gemini text-to-speech models, and it frames them as tools for a creative studio instead of simple preset voice boxes. According to the company, Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS allow creators to construct entirely new voices from nothing, deliver direct line-by-line performances, and even copy a person’s voice from a 30-second sample.
Two Models, One Creative Studio
The two models serve different purposes. Gemini 3.8 Flash TTS is built for deep creative direction and character design. You can create entirely new voices using natural language prompts, direct every performance line by line, and shift dialects mid-sentence. It’s aimed at gaming, immersive audiobooks, podcasts, and interactive media.
The high-volume edition of Gemini Flash-Lite TTS is built for dubbing, audio content creation, and expressive voice agents. Fine-grained control over tone, pacing, and expressive nuance is available on both models. 3.8
The Gemini Audio family is expanding with these new additions, and it now counts 3.5 Live Translate, 3.5 Transcribe, 3.8 Live, and 3.8 Live Extended Thinking among its members.
Building Voices From Scratch
Gemini 3.8 Flash TTS offers to take you from nothing to a working voice by letting you tailor role, accent, and voice traits across over 100 languages and dialects through natural language prompts. It includes a dramatic example of a fire-breathing dragon and a narrator with a clear regional cadence.
A large selection of production-ready voices is part of what the company provides. More than 2,000 voices are included, ranging from Mexican Spanish to Quebec French to Scots English. A consistent vocal profile can be reproduced from just a 30-second audio clip, complete with built-in consent verification, SynthID watermarking, and C2PA credentials designed to protect both developers and the voice artists they work with.
Consistency across ongoing projects is maintained through save and scale features, which preserve custom voices. A forthcoming voice remixing tool will allow selection of a voice from the library and adjustment of timbre, pitch, pace, and accent using prompts such as “add subtle Southern US accent” or “soften the delivery.”.
Line-by-Line Direction
Both TTS models allow you to control delivery after you pick your voices. You can compose your own stage directions or let Gemini handle things with natural script prompts instead. The models keep high voice quality, natural pacing, and character timbre across hours of continuous audio, making them well suited for podcasts and audiobooks.
A single script can be used to stage conversations between two speakers, with their voices kept apart from one another. The text includes scripted bursts of speech along with background reactions that add natural sound to the talk, including laughs, sighs, gasps, and listening responses such as |mhm| or |yeah|.
Benchmarks And Blind Tests
Gemini 3.8 Flash TTS secured the #1 overall spot on Hume AI’s Voice Design Benchmark, scoring 71.4. It also led in accent modeling, scoring 60.8. The model holds the #1 and #2 spots respectively on Hume AI’s Overall Quality Index, with major improvements on long-form content and dual-speaker screenplay control compared to Gemini 3.1 Flash TTS.
Gemini 3.8 Flash and Flash-Lite TTS placed first in human preference tests on Voice Arena across a number of major world languages. Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic (MSA), Mexican Spanish, and Hindi all featured among the key languages where the two models led their rivals. Together, the models support more than 100 languages.
| Model | Focus | Key Feature |
|---|---|---|
| Gemini 3.8 Flash TTS | Creative direction | New voices from scratch, line-by-line direction |
| Gemini 3.8 Flash-Lite TTS | High-volume scale | Dubbing, voice agents, fine-grained control |
Trust, Consent, And Watermarking
These tools were constructed with strict safeguards in place by Google. To produce a voice, a user must first present a verbal consent recording from the voice’s owner that corresponds to the reference speaker. Each audio clip is marked with SynthID, an imperceptible watermark woven directly into the audio output so that AI-generated speech can be detected.
Readers are directed by the company to the model card for further information on its approach to safety and responsibility.
Try It In Google AI Studio
In Google AI Studio, developers can explore fresh speech generation features through a voice design workspace. They can start with an entirely new vocal identity or reproduce their own voice, then move those creations into a dual-speaker screenplay editor where line-by-line delivery is directed.
What We Make Of It
Generative audio has taken a notable leap forward. Creating entirely new voices from nothing, directing performances with fine-grained control, and copying a person’s voice with consent verification all push text-to-speech past fixed presets toward a creative toolkit instead.
Real gains over past models appear in the benchmarks and blind tests, pointing to a clear push by the company toward audio that delivers greater expressiveness at scale.
A practical safeguard against copying a voice without permission is the consent verification requirement, and it’s worth noting because doing so could otherwise be a serious legal problem.
Some things remain unclear. The dual-speaker setup gets a mention but no further details. The company has also kept quiet on whether these models will join the existing Gemini TTS lineup.
Still, this is a strong update for the Gemini Audio family. The new models give creators more freedom to build custom audio experiences, and the company’s commitment to consent and watermarking suggests it’s thinking seriously about the ethical implications of this technology.
No longer a mere tool, text-to-speech has become a creative medium, and the announcement makes clear that Google sees itself taking the lead in that change.
Source material: “Gemini 3.8 text-to-speech,” Google.
Get the Notebook.
The day's best stories and every fresh verdict, in plain English, in your inbox by seven. One email a day, no more.

