Gemini 3.5 Transcribe icon

Gemini 3.5 Transcribe

Gemini 3.5 Transcribe is Google’s speech-to-text model for real-time and recorded audio transcription. It helps developers and product users turn spoken language into clean, formatted text with support for multilingual audio, timestamps, and speaker attribution.

Gemini 3.5 Transcribe

Overview

Gemini 3.5 Transcribe is Google’s speech-to-text model for intelligent transcription. The announcement positions it as a more precise option for real-time voice interactions and for converting recorded audio into clean, formatted text.

The model is designed to handle common transcription problems such as background noise, jargon, disfluencies, and mid-sentence corrections. It is available through the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform, and it also powers voice features in Google products such as the Gemini app on macOS and Rambler on Android.

Key capabilities

Intelligent real-time transcription

The model is built to transcribe speech into polished text even when audio includes noise, jargon, or self-corrections.

Low-latency streaming via Live API

The Live API supports continuous bidirectional streaming with sub-second latency for interactive voice applications.

Pre-recorded audio processing

The Interactions API handles recorded audio, meetings, and call logs, and can return speaker attribution plus word-level timestamps.

Multilingual language detection

The model automatically detects and transcribes more than 85 languages, including regional accents and dialects, and can handle live language switches.

Custom vocabulary support

It recognizes specialized terms and unique spellings through custom vocabulary supplied by the user.

Voice-enabled function calling

The model can delegate tasks such as image generation and file analysis through function calls, and this is currently available in the Gemini macOS app.

Practical use cases

  • Interactive voice applications

    Build voice agents and other interactive apps that need continuous speech-to-text with low latency and bidirectional streaming.

  • Meeting and call transcription

    Transcribe meetings, call recordings, and audio logs, then use speaker attribution and timestamps to review what was said and by whom.

  • Live captioning

    Create real-time captioning tools that need transcription output while audio is still being captured.

  • Voice-driven productivity

    Use voice commands in the Gemini app on macOS or Android workflows to dictate, edit, summarize, or clean up speech into polished text.

  • Post-call analytics

    Process recorded audio for post-call analytics pipelines where accurate formatted text is needed for downstream analysis.

Pros and Cons

Pros

  • Supports both real-time streaming and pre-recorded transcription workflows.
  • Includes speaker attribution and word-level timestamps for recorded audio.
  • Automatically detects and transcribes more than 85 languages.
  • Handles background noise, jargon, filler words, and self-corrections.
  • Available to developers through Google AI Studio and the Gemini Enterprise Agent Platform.

Cons

  • Pricing is not published on the source page.
  • Some consumer features are only partially available, such as Rambler on Android in select countries and languages, Chrome coming soon, and 3+ speaker support marked experimental.

FAQ

How do developers access Gemini 3.5 Transcribe?

Developers can access Gemini 3.5 Transcribe in the Gemini API through Google AI Studio and the Gemini Enterprise Agent Platform. The announcement also says it is in public preview for developers.

What kinds of transcription workflows does it support?

The page describes two API paths: the Live API for real-time streaming with sub-second latency, and the Interactions API for pre-recorded audio with speaker attribution and word-level timestamps.

What transcription challenges is it designed to handle?

It is designed to handle background noise, complex jargon, disfluency cleanup, custom vocabulary, and more than 85 languages, including regional accents and dialects.

Where is Gemini 3.5 Transcribe available in Google products?

The announcement says the model is available to everyone in the Gemini app on macOS in English, in Rambler on Android in select countries and languages, and is coming soon to Chrome.

Is pricing published for Gemini 3.5 Transcribe?

The source does not provide public pricing details. It only states that the model is in public preview across developer and enterprise access points.

Quick Facts

Category
Speech-to-text / transcription
Primary users
Developers, enterprises, and consumer app users
Access
Gemini API in Google AI Studio; Gemini Enterprise Agent Platform; Gemini app on macOS; Rambler on Android
API modes
Live API for real-time streaming; Interactions API for pre-recorded audio
Languages
85+ languages with regional accents and dialects
Source domain
blog.google