Gemini 3.8 Text-to-Speech icon

Gemini 3.8 Text-to-Speech

Gemini 3.8 Text-to-Speech is Google’s pair of expressive audio models for turning scripts into customizable spoken audio. Creators, developers and enterprises can design voices, direct delivery line by line, and generate dialogue for content, dubbing and voice-agent applications.

Gemini 3.8 Text-to-Speech

Overview

Gemini 3.8 Text-to-Speech is a pair of expressive audio-generation models from Google: Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS. They turn text and scripts into directed speech, moving beyond fixed voice presets with prompt-based voice design, detailed performance instructions and support for multi-speaker scenes.

Flash TTS is intended for creative voice and character design, while Flash-Lite TTS is positioned for high-volume dubbing, audio creation and expressive voice agents. The source lists Google AI Studio, the Gemini API, Gemini Enterprise, Gemini Notebook and Google Vids as product contexts for these models, although it does not provide detailed access, pricing or quota information.

Key features

Prompt-based voice design

Gemini 3.8 Flash TTS can create bespoke voices by prompting for a role, accent and other vocal characteristics across more than 100 languages and dialects.

Large voice library

The source describes a library of more than 2,000 production-ready voices, including regional varieties such as Mexican Spanish, Quebec French and Scots English.

Line-by-line performance control

Users can direct each line with natural-language stage directions for pacing, emotion, dialect shifts, delivery style and conversational reactions.

Two-speaker scene staging

Scripts can stage two speakers in a single scene while preserving distinct voices and natural conversational turn-taking.

Long-form generation

The models support extended audio generation for podcasts and audiobooks, with natural pacing and character timbre intended to remain consistent over hours of audio.

Voice replication and provenance tools

Voice replication can use a 30-second sample of a voice the user owns or has rights to use. The source also cites consent verification, SynthID watermarking and C2PA credentials for this workflow.

Use cases

  • Character voices and interactive media

    Design a recurring character or narrator with a prompted vocal identity, then direct individual lines for dramatic scenes, games or immersive storytelling.

  • Audiobooks and podcasts

    Produce long-form spoken content with controlled pacing and consistent character timbre for audiobooks and podcasts.

  • Dubbing and content localization

    Create localized or high-volume spoken content using the Flash-Lite model’s positioning for dubbing and expressive audio production.

  • Voice agents

    Build voice agents that use tone, pacing, backchanneling and conversational reactions to make spoken interactions more expressive.

  • Two-person dialogue scenes

    Write a multi-turn script with two distinct speakers and control their handoffs and reactions for narrative or podcast-style scenes.

Pros and Cons

Pros

  • Supports custom voice creation from natural-language prompts, including role, accent and vocal characteristics.
  • Provides detailed script control for pacing, emotion, nonverbal cues and conversational backchanneling.
  • Offers two-speaker scene staging and long-form generation for extended narrative audio.
  • Includes a broad voice library and language coverage spanning more than 100 languages and dialects.
  • Describes consent verification, SynthID watermarking and C2PA credentials for voice-replication workflows.

Cons

  • The source does not document pricing, quotas, output formats, latency or detailed availability requirements.
  • Voice remixing—fine-tuning an existing library voice’s timbre, pitch, pace or accent—is listed as coming soon rather than generally available.
  • Voice replication should be used only with a sample the user owns or has permission to use; the source does not provide further operational or legal guidance.

FAQ

What is Gemini 3.8 Text-to-Speech?

The product includes two models: Gemini 3.8 Flash TTS for deeper creative direction and custom character voices, and Gemini 3.8 Flash-Lite TTS for high-volume, cost-efficient generation. Both support line-by-line performance direction.

Where can the models be used?

The source lists Google AI Studio, the Gemini API, Gemini Enterprise, Gemini Notebook and Google Vids as places where the models can be used. Specific availability, account requirements and regional restrictions are not provided in the source.

Can users create or replicate a voice?

Yes. Flash TTS can design voices from natural-language prompts, and the source describes voice replication from a 30-second sample of a voice that the user owns or has permission to use. Consent verification, SynthID watermarking and C2PA credentials are described as part of the replication workflow.

How can generated speech be directed?

Users can direct delivery line by line with stage directions and script cues, including pacing, emotion, dialect changes, vocal bursts and backchanneling. The models also support native two-speaker scene staging and long-form generation with voice consistency across extended audio.

Quick Facts

Product type
Text-to-speech and expressive audio generation models
Models
Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS
Primary users
Creators, developers and enterprises
Language coverage
More than 100 languages and dialects
Listed product contexts
Google AI Studio, Gemini API, Gemini Enterprise, Gemini Notebook and Google Vids
Source domain
blog.google
Gemini 3.8 Text-to-Speech - AI Tool, Features, Use Cases & Alternatives | UStack