Native code-switching across 18 languages
The launch article says the model handles code-switched speech natively across 18 languages, so transcripts can follow the language actually spoken without a separate language pass.
Universal-3.5 Pro is AssemblyAI’s async speech-to-text model for recorded audio and video. It adds native code-switching, speaker diarization, and contextual prompting, with a base price of $0.21 per hour of audio submitted.
Universal-3.5 Pro is AssemblyAI’s flagship asynchronous speech-to-text model for recorded audio and video. It is part of the Pre-recorded Speech-to-Text API and is positioned as the highest-accuracy model for supported languages and difficult real-world audio.
The launch article highlights three main capabilities: native code-switching across 18 languages, improved speaker diarization, and contextual prompting. The pricing page lists the base rate at $0.21 per hour of audio submitted, with optional add-ons for capabilities such as diarization, medical mode, keyterms prompting, and general prompting.
In practice, the model is aimed at transcription workloads where speech is messy, multilingual, or speaker-heavy. That includes meetings, contact center recordings, captions, searchable archives, and other recorded workflows where accuracy and speaker attribution matter more than live latency.
The launch article says the model handles code-switched speech natively across 18 languages, so transcripts can follow the language actually spoken without a separate language pass.
AssemblyAI describes this release as its most accurate speaker diarization yet, with the model designed to keep speaker attribution aligned with the transcript even in crosstalk and noisy conversations.
Contextual prompting lets you provide domain knowledge or prior context to influence the transcription output, which can improve accuracy on specialized terms or ambiguous audio.
The pricing reference lists keyterms prompting for Universal-3.5 Pro with support for up to 1,000 terms, giving teams a way to bias recognition toward project-specific vocabulary.
Universal-3.5 Pro is the async pre-recorded model in AssemblyAI’s speech-to-text lineup and is priced at $0.21 per hour of audio submitted.
The pricing page shows optional add-ons for this model, including medical mode and speaker diarization, which can be combined with the base transcription rate.
Use Universal-3.5 Pro when a recording contains multiple languages in the same conversation and you want the transcript to preserve the language spoken at each point.
Use the model for meeting notes or internal discussions where short turns, interruptions, and overlapping voices make speaker attribution difficult.
Use it for contact center calls or other customer interactions where accurate speaker labeling affects downstream analysis and QA.
Use contextual prompting or keyterms prompting when transcribing specialized discussions that include project names, domain terms, or other vocabulary that may be misheard.
Use the async API for archives, captions, or dataset preparation when the audio is already recorded and you can trade latency for transcription quality.
Universal-3.5 Pro is AssemblyAI’s async pre-recorded speech-to-text model. It is designed for recorded audio and video rather than live streaming transcription.
The launch article says it supports native code-switching across 18 languages. The product pages identify it as the highest-accuracy pre-recorded model for supported languages.
The launch article highlights contextual prompting, and the pricing reference shows general prompting is available as a paid async add-on for Universal-3.5 Pro.
The pricing page lists async speaker diarization for Universal-3.5 Pro as a paid add-on, with standard and experimental variants. The launch article positions diarization as one of the model’s core improvements.
The pricing page lists Universal-3.5 Pro at $0.21 per hour for pre-recorded audio submitted. Add-ons such as diarization, medical mode, keyterms prompting, and general prompting are priced separately.
QuickQuill is a macOS dictation and transcription app that runs locally on the device. It helps users record meetings, transcribe audio, generate summaries, and export notes without using a cloud service.
Speech to Text Converter is a browser-based transcription tool for live dictation and uploaded audio or video files. It offers a free tier for short tasks and a Pro plan for unlimited transcription, AI summaries, translation, speaker identification, and advanced exports.
An OpenAI API guide for choosing the right speech architecture for live audio, translation, transcription, speech generation, and audio-capable chat. It helps developers map each speech application to the appropriate session type, endpoint, and connection method.
Pewbeam is a church presentation app that listens to sermons, detects Bible verse references in real time, and displays the matching passage on screen. It is built for pastors, projection teams, and church media volunteers who want to reduce manual slide control during live services.
Dictato is a Mac dictation app that transcribes speech into text in any app using an on-device, offline workflow. It supports multiple transcription engines, optional cleanup and translation, and a one-time purchase license.
Voicenotes is an AI note-taking and meeting recording app that transcribes conversations, generates summaries and action items, and makes past recordings searchable with Ask AI. It also supports voice dictation, imported audio on Pro, and multilingual transcription.