Google releases Gemini 3.8 Flash TTS and Flash-Lite TTS through its cloud platform
Google has made Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS available on its cloud platform, two speech generation models that took the top two spots on a Hume AI audio quality benchmark.
The models have highly similar application programming interfaces, which Google says makes it relatively simple for developers to use them side by side. Flash-Lite TTS is tuned for cost efficiency and inference speed, while Flash TTS produces better audio at a higher price. Google expects customers to use the higher-end model for tasks such as audiobook production.
The two models differ in language coverage. Flash TTS can generate speech in 130 languages at launch, against 101 for Flash-Lite TTS.
Both give developers access to a library of more than 2,000 prepackaged voices and allow custom voices to be created with natural language prompts. Parameters such as a speaker's vocal timbre, accent and pacing can be adjusted.
A second customization route generates a synthetic voice from a 30-second audio sample. Google requires developers to obtain the speaker's consent before generating a voice replica. The company said it plans to add a third option later that would create a new voice by modifying one of the prepackaged choices.
Google also allows changes to how a voice delivers a script. Developers can attach oratory cues to individual lines that the models read aloud, producing audio elements such as non-lexical vocalizations and shifts in pacing.
Generated speech carries an audio watermark embedded with a technology called SynthID. The watermark is inaudible to listeners but can be detected by AI detection tools. Each generated audio file also carries a C2PA record, which states when the file was created, whether it has been modified since and related details.
On the Hume AI audio quality benchmark, Flash TTS placed first and Flash-Lite TTS second. Google said the models also outperformed several competing systems on multiple language-specific versions of Voice Arena, a benchmark that measures text-to-speech output quality based on human feedback.
"These models enable creators, developers, and enterprises to create richer, more expressive audio experiences, while enabling improved user experiences in products like Gemini Notebook and Google Vids," Google staffers Leland Rechis and Alan Cowen wrote in a blog post.
The two models belong to a wider line of audio processing algorithms from Google, which has previously released models optimized for voice agents, transcription and translation.