Prompt-based voice design
Gemini 3.8 Flash TTS can create bespoke voices by prompting for a role, accent and other vocal characteristics across more than 100 languages and dialects.
Gemini 3.8 Text-to-Speech is Google’s pair of expressive audio models for turning scripts into customizable spoken audio. Creators, developers and enterprises can design voices, direct delivery line by line, and generate dialogue for content, dubbing and voice-agent applications.
Gemini 3.8 Text-to-Speech is a pair of expressive audio-generation models from Google: Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS. They turn text and scripts into directed speech, moving beyond fixed voice presets with prompt-based voice design, detailed performance instructions and support for multi-speaker scenes.
Flash TTS is intended for creative voice and character design, while Flash-Lite TTS is positioned for high-volume dubbing, audio creation and expressive voice agents. The source lists Google AI Studio, the Gemini API, Gemini Enterprise, Gemini Notebook and Google Vids as product contexts for these models, although it does not provide detailed access, pricing or quota information.
Gemini 3.8 Flash TTS can create bespoke voices by prompting for a role, accent and other vocal characteristics across more than 100 languages and dialects.
The source describes a library of more than 2,000 production-ready voices, including regional varieties such as Mexican Spanish, Quebec French and Scots English.
Users can direct each line with natural-language stage directions for pacing, emotion, dialect shifts, delivery style and conversational reactions.
Scripts can stage two speakers in a single scene while preserving distinct voices and natural conversational turn-taking.
The models support extended audio generation for podcasts and audiobooks, with natural pacing and character timbre intended to remain consistent over hours of audio.
Voice replication can use a 30-second sample of a voice the user owns or has rights to use. The source also cites consent verification, SynthID watermarking and C2PA credentials for this workflow.
Design a recurring character or narrator with a prompted vocal identity, then direct individual lines for dramatic scenes, games or immersive storytelling.
Produce long-form spoken content with controlled pacing and consistent character timbre for audiobooks and podcasts.
Create localized or high-volume spoken content using the Flash-Lite model’s positioning for dubbing and expressive audio production.
Build voice agents that use tone, pacing, backchanneling and conversational reactions to make spoken interactions more expressive.
Write a multi-turn script with two distinct speakers and control their handoffs and reactions for narrative or podcast-style scenes.
The product includes two models: Gemini 3.8 Flash TTS for deeper creative direction and custom character voices, and Gemini 3.8 Flash-Lite TTS for high-volume, cost-efficient generation. Both support line-by-line performance direction.
The source lists Google AI Studio, the Gemini API, Gemini Enterprise, Gemini Notebook and Google Vids as places where the models can be used. Specific availability, account requirements and regional restrictions are not provided in the source.
Yes. Flash TTS can design voices from natural-language prompts, and the source describes voice replication from a 30-second sample of a voice that the user owns or has permission to use. Consent verification, SynthID watermarking and C2PA credentials are described as part of the replication workflow.
Users can direct delivery line by line with stage directions and script cues, including pacing, emotion, dialect changes, vocal bursts and backchanneling. The models also support native two-speaker scene staging and long-form generation with voice consistency across extended audio.
Fylom is an AI podcast generator that turns a topic or question into a researched narrated episode. It supports typed or spoken prompts, voice selection, and recurring show-style listening.
Gemini 3.1 Flash TTS is Google’s preview text-to-speech model for generating expressive AI speech with fine-grained control over style and delivery. It is available across the Gemini API, Google AI Studio, Vertex AI, and Google Vids.
蓝藻AI是一款在线AI配音与语音合成产品,可将文字转成语音,并支持自助声音克隆。页面信息显示它面向短视频、有声书等需要配音的内容场景。
Ondoku 是一款基于浏览器的文字转语音软件,可将文本转换为可下载的 .mp3 语音,并提供免费额度与付费方案。它支持多语言朗读、图片朗读以及按规则商用。
Typecast is an online AI voice generator that turns text into life-like speech with emotional delivery and a selection of hyper-realistic voices. It is a browser-based tool for creating spoken audio from written content.
Noiz AI is an AI text-to-speech, voice cloning, and voice design tool for creating lifelike speech from text. It also lets users shape voice delivery, including emotion, within the same workflow.