Intelligent real-time transcription
The model is built to transcribe speech into polished text even when audio includes noise, jargon, or self-corrections.
Gemini 3.5 Transcribe is Google’s speech-to-text model for real-time and recorded audio transcription. It helps developers and product users turn spoken language into clean, formatted text with support for multilingual audio, timestamps, and speaker attribution.
Gemini 3.5 Transcribe is Google’s speech-to-text model for intelligent transcription. The announcement positions it as a more precise option for real-time voice interactions and for converting recorded audio into clean, formatted text.
The model is designed to handle common transcription problems such as background noise, jargon, disfluencies, and mid-sentence corrections. It is available through the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform, and it also powers voice features in Google products such as the Gemini app on macOS and Rambler on Android.
The model is built to transcribe speech into polished text even when audio includes noise, jargon, or self-corrections.
The Live API supports continuous bidirectional streaming with sub-second latency for interactive voice applications.
The Interactions API handles recorded audio, meetings, and call logs, and can return speaker attribution plus word-level timestamps.
The model automatically detects and transcribes more than 85 languages, including regional accents and dialects, and can handle live language switches.
It recognizes specialized terms and unique spellings through custom vocabulary supplied by the user.
The model can delegate tasks such as image generation and file analysis through function calls, and this is currently available in the Gemini macOS app.
Build voice agents and other interactive apps that need continuous speech-to-text with low latency and bidirectional streaming.
Transcribe meetings, call recordings, and audio logs, then use speaker attribution and timestamps to review what was said and by whom.
Create real-time captioning tools that need transcription output while audio is still being captured.
Use voice commands in the Gemini app on macOS or Android workflows to dictate, edit, summarize, or clean up speech into polished text.
Process recorded audio for post-call analytics pipelines where accurate formatted text is needed for downstream analysis.
Developers can access Gemini 3.5 Transcribe in the Gemini API through Google AI Studio and the Gemini Enterprise Agent Platform. The announcement also says it is in public preview for developers.
The page describes two API paths: the Live API for real-time streaming with sub-second latency, and the Interactions API for pre-recorded audio with speaker attribution and word-level timestamps.
It is designed to handle background noise, complex jargon, disfluency cleanup, custom vocabulary, and more than 85 languages, including regional accents and dialects.
The announcement says the model is available to everyone in the Gemini app on macOS in English, in Rambler on Android in select countries and languages, and is coming soon to Chrome.
The source does not provide public pricing details. It only states that the model is in public preview across developer and enterprise access points.
SpeakoFlow is a free, open-source desktop voice-to-text app for Windows, macOS, and Linux. It supports dictation into any app, voice-driven writing with Flow, cleanup, translation, and a screen-aware assistant.
QuickQuill is a macOS dictation and transcription app that runs locally on the device. It helps users record meetings, transcribe audio, generate summaries, and export notes without using a cloud service.
Zen Whisper is a local-first dictation and transcription app for Apple Silicon Macs running macOS 14 or later. It lets users dictate into most Mac text fields, transcribe media or supported public links, and work across 110 spoken language options.
Speech to Text Converter is a browser-based transcription tool for live dictation and uploaded audio or video files. It offers a free tier for short tasks and a Pro plan for unlimited transcription, AI summaries, translation, speaker identification, and advanced exports.
Sanota is an app that turns spoken memories, reflections, and interviews into clear written stories. It supports personal storytelling, family history, and shared memories, with guided prompts and subscription pricing.
Carbon Voice is an asynchronous voice messaging app for teams and individuals, with transcripts, AI catch-up, and cross-device access. It helps people and agents communicate without needing a live call.