Speaker diarization
Voxtral Mini Transcribe V2 generates speaker-labeled transcripts with precise start and end times, which is useful when you need to know who said what and when.
Voxtral Transcribe 2 by Mistral AI offers batch and live speech-to-text with diarization, timestamps, multilingual support, and an audio playground in Mistral Studio.
Voxtral Transcribe 2 is Mistral AI’s speech-to-text offering, introduced as two next-generation models for transcription: Voxtral Mini Transcribe V2 for batch transcription and Voxtral Realtime for live applications. The launch focuses on transcription quality, speaker diarization, low latency, and language coverage rather than on a broader conversation or agent platform.
The product page says Voxtral Mini Transcribe V2 provides state-of-the-art transcription with diarization, context biasing, and word-level timestamps in 13 languages, while Voxtral Realtime is designed for streaming audio with latency configurable down to sub-200ms. Mistral also says the Realtime model is open-weights under Apache 2.0, and that an audio playground in Mistral Studio lets users test transcription with diarization and timestamps before building against the API.
The source positions Voxtral for workflows such as meeting transcription, voice agents, contact center automation, media subtitling, and compliance documentation. Pricing details in the article indicate API access for both models, with Mini Transcribe V2 at $0.003 per minute and Realtime at $0.006 per minute, plus an open-weights release for Realtime on Hugging Face.
Voxtral Mini Transcribe V2 generates speaker-labeled transcripts with precise start and end times, which is useful when you need to know who said what and when.
You can provide up to 100 words or phrases to bias the model toward names, technical terms, and other vocabulary that standard transcription systems may miss.
The model can return timestamps for each word, supporting subtitle creation, searchable archives, and time-aligned content workflows.
Both models support 13 languages: English, Chinese, Hindi, Spanish, Arabic, French, Portuguese, Russian, German, Japanese, Korean, Italian, and Dutch.
Voxtral Realtime is built for live audio and supports latency configurable down to sub-200ms, while Mini Transcribe V2 is positioned for batch transcription.
The product description highlights an audio playground in Mistral Studio for immediate testing with diarization, timestamps, and audio file uploads.
Transcribe recurring meetings with speaker labels and timestamps so teams can review decisions, assignments, and discussion flow after the call.
Power conversational agents and assistant experiences that need transcription latency low enough to keep voice interactions responsive.
Process customer support or sales calls as they happen, using diarization to separate agent and customer speech for later analysis or CRM entry.
Generate live or near-live subtitles for multilingual media, where low latency and word-level timing help align speech with on-screen captions.
Record regulated or sensitive conversations with diarization and timestamps to create clearer audit trails for review and documentation.
Voxtral Transcribe 2 is a speech-to-text product family with two model options: Voxtral Mini Transcribe V2 for batch transcription and Voxtral Realtime for live applications. The article also mentions an audio playground in Mistral Studio for testing transcription directly.
The source describes Voxtral Mini Transcribe V2 as a batch transcription model, and Voxtral Realtime as a streaming model designed for live applications where low latency matters. The article does not provide a full API or SDK workflow beyond these product names and the Mistral Studio playground.
According to the source, the audio playground in Mistral Studio supports uploading up to 10 audio files, toggling diarization, choosing timestamp granularity, and adding context bias terms. It accepts .mp3, .wav, .m4a, .flac, and .ogg files up to 1GB each.
The article states that Voxtral Mini Transcribe V2 is available via API at $0.003 per minute, while Voxtral Realtime is available via API at $0.006 per minute and as open weights on Hugging Face. The pricing page also confirms that Mistral offers API usage and a Studio dashboard, but it does not add Voxtral-specific packaging details.
The source says Voxtral Realtime is open-weights under the Apache 2.0 license and can be deployed on edge devices. It also says both models support secure on-premise or private cloud setups for GDPR- and HIPAA-compliant deployments, but the article does not provide implementation steps.
QuickQuill is a local-first macOS dictation and transcription app to record meetings, summarize audio, and export notes without the cloud.
Speech to Text Converter is a browser-based transcription tool for live dictation and uploaded audio or video files. Free for short tasks, Pro offers unlimited transcription, AI summaries, translation, speaker ID, and advanced exports.
Dictato is a Mac dictation app that transcribes speech to text in any app with an offline, on-device workflow. Includes cleanup, translation, and one-time purchase.
Sanota turns spoken memories and interviews into clear written stories for personal storytelling, family history and shared memories, with guided prompts and subscriptions.
Carbon Voice is an async voice messaging app for teams and individuals, with transcripts, AI catch-up, and cross-device access.
OpenAI API guide for choosing the right speech architecture for live audio, translation, transcription, speech generation, and audio-capable chat.