topping speech generation models
Google has released two new text-to-speech models, Gemini 3.8 Flash and Flash-Lite, offering varying levels of quality and language support for developers. The models include advanced customization features and security measures like audio watermarking and C2PA records to track AI-generated content.
Why it matters
The integration of watermarking and provenance records represents a significant step in the industry's effort to combat deepfakes and verify the authenticity of AI-generated audio.
Google LLC today made two new text-to-speech models, Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, available through its cloud platform.
The algorithms have highly similar application programming interfaces, which makes using them side-by-side relatively simple for developers. Flash-Lite TTS is optimized for cost-cost efficiency and inference speed. Flash TTS offers better audio quality for a higher price. Google envisions customers using it for tasks such as creating audiobooks.
There are also other differences between the models. Most notably, Flash TTS can generate speech in 130 languages on launch while Flash-Lite TTS supports 101.
Both models offer access to a library of more than 2,000 prepackaged voices. Developers can create custom voices with natural language prompts. Google makes it possible to customize parameters such an AI speaker’s vocal timbre, accent and pacing.
Also covering this story
One other newsroom covered this event. We read that version too.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in