Improved speech quality
The model is presented as Google’s most natural and expressive text-to-speech model to date, with improved speech quality and controllability.
Gemini 3.1 Flash TTS is Google’s preview text-to-speech model for generating expressive AI speech with fine-grained control over style and delivery. It is available across the Gemini API, Google AI Studio, Vertex AI, and Google Vids.
Gemini 3.1 Flash TTS is Google’s text-to-speech model for generating expressive AI speech with tighter control over how audio sounds. The launch announcement emphasizes improved naturalness, clearer pacing control, and new audio tags that let developers direct vocal style and delivery through natural-language instructions.
The model is rolling out in preview for developers through the Gemini API and Google AI Studio, for enterprises through Vertex AI, and for Workspace users through Google Vids. It supports 70+ languages, native multi-speaker dialogue, and SynthID watermarking for every generated audio output.
The model is presented as Google’s most natural and expressive text-to-speech model to date, with improved speech quality and controllability.
Audio tags let users direct vocal style, pace, delivery, tone, and accent with natural-language instructions embedded in the text input.
Google AI Studio adds configurable controls for scene direction, speaker-level specificity, and inline tags, helping developers shape multi-turn performances.
Developers can export the exact voice parameters from Google AI Studio as Gemini API code for consistent reuse across projects and platforms.
The model supports native multi-speaker dialogue and over 70 languages, which makes it suitable for localized and conversational speech experiences.
All generated audio is watermarked with SynthID to support detection of AI-generated content.
Build applications that need synthesized speech with controlled delivery, such as character voices, narrated experiences, or interactive assistants.
Prototype voice experiences in Google AI Studio, refine pacing and tone with tags and notes, and export the resulting settings into Gemini API code.
Create localized speech experiences for audiences across multiple languages while keeping style and accent control consistent.
Use the model in Google Vids when you need AI-generated speech for Workspace media workflows.
Generate audio with built-in SynthID watermarking when you need detectable AI-generated speech for safer distribution.
It is rolling out for developers in preview through the Gemini API and Google AI Studio, for enterprises in preview on Vertex AI, and for Workspace users through Google Vids.
The announcement says it supports 70+ languages and includes native multi-speaker dialogue.
Developers can use audio tags, Audio Profiles, Director’s Notes, and inline tags in Google AI Studio to steer vocal style, pacing, tone, accent, and speaker delivery, then export the same parameters as Gemini API code.
All audio generated by Gemini 3.1 Flash TTS is watermarked with SynthID, which is described as an imperceptible watermark for detecting AI-generated audio.
The source does not provide pricing details on the product page, and the pricing page linked in the research set is a 404.
蓝藻AI是一款在线AI配音与语音合成产品,可将文字转成语音,并支持自助声音克隆。页面信息显示它面向短视频、有声书等需要配音的内容场景。
Ondoku 是一款基于浏览器的文字转语音软件,可将文本转换为可下载的 .mp3 语音,并提供免费额度与付费方案。它支持多语言朗读、图片朗读以及按规则商用。
Typecast is an online AI voice generator that turns text into life-like speech with emotional delivery and a selection of hyper-realistic voices. It is a browser-based tool for creating spoken audio from written content.
Noiz AI is an AI text-to-speech, voice cloning, and voice design tool for creating lifelike speech from text. It also lets users shape voice delivery, including emotion, within the same workflow.
魔音工坊 (Moying Gongfang) es una plataforma inteligente de texto a voz (TTS) en línea que convierte texto escrito en locuciones de alta calidad utilizando voces humanas realistas con diversos acentos.
TADA is Hume AI’s open-source speech-language model for generating speech with one-to-one text-acoustic alignment. It is aimed at developers and researchers building faster, more reliable voice systems, including on-device and long-form speech applications.