Native code-switching across 18 languages
The launch article says the model handles code-switched speech natively across 18 languages, so transcripts can follow the language actually spoken without a separate language pass.
Universal-3.5 Pro is AssemblyAI’s async speech-to-text model for recorded audio and video. Native code-switching, diarization, contextual prompting; $0.21/hour.
Universal-3.5 Pro is AssemblyAI’s flagship asynchronous speech-to-text model for recorded audio and video. It is part of the Pre-recorded Speech-to-Text API and is positioned as the highest-accuracy model for supported languages and difficult real-world audio.
The launch article highlights three main capabilities: native code-switching across 18 languages, improved speaker diarization, and contextual prompting. The pricing page lists the base rate at $0.21 per hour of audio submitted, with optional add-ons for capabilities such as diarization, medical mode, keyterms prompting, and general prompting.
In practice, the model is aimed at transcription workloads where speech is messy, multilingual, or speaker-heavy. That includes meetings, contact center recordings, captions, searchable archives, and other recorded workflows where accuracy and speaker attribution matter more than live latency.
The launch article says the model handles code-switched speech natively across 18 languages, so transcripts can follow the language actually spoken without a separate language pass.
AssemblyAI describes this release as its most accurate speaker diarization yet, with the model designed to keep speaker attribution aligned with the transcript even in crosstalk and noisy conversations.
Contextual prompting lets you provide domain knowledge or prior context to influence the transcription output, which can improve accuracy on specialized terms or ambiguous audio.
The pricing reference lists keyterms prompting for Universal-3.5 Pro with support for up to 1,000 terms, giving teams a way to bias recognition toward project-specific vocabulary.
Universal-3.5 Pro is the async pre-recorded model in AssemblyAI’s speech-to-text lineup and is priced at $0.21 per hour of audio submitted.
The pricing page shows optional add-ons for this model, including medical mode and speaker diarization, which can be combined with the base transcription rate.
Use Universal-3.5 Pro when a recording contains multiple languages in the same conversation and you want the transcript to preserve the language spoken at each point.
Use the model for meeting notes or internal discussions where short turns, interruptions, and overlapping voices make speaker attribution difficult.
Use it for contact center calls or other customer interactions where accurate speaker labeling affects downstream analysis and QA.
Use contextual prompting or keyterms prompting when transcribing specialized discussions that include project names, domain terms, or other vocabulary that may be misheard.
Use the async API for archives, captions, or dataset preparation when the audio is already recorded and you can trade latency for transcription quality.
Universal-3.5 Pro is AssemblyAI’s async pre-recorded speech-to-text model. It is designed for recorded audio and video rather than live streaming transcription.
The launch article says it supports native code-switching across 18 languages. The product pages identify it as the highest-accuracy pre-recorded model for supported languages.
The launch article highlights contextual prompting, and the pricing reference shows general prompting is available as a paid async add-on for Universal-3.5 Pro.
The pricing page lists async speaker diarization for Universal-3.5 Pro as a paid add-on, with standard and experimental variants. The launch article positions diarization as one of the model’s core improvements.
The pricing page lists Universal-3.5 Pro at $0.21 per hour for pre-recorded audio submitted. Add-ons such as diarization, medical mode, keyterms prompting, and general prompting are priced separately.
QuickQuill is a local-first macOS dictation and transcription app to record meetings, summarize audio, and export notes without the cloud.
Speech to Text Converter is a browser-based transcription tool for live dictation and uploaded audio or video files. Free for short tasks, Pro offers unlimited transcription, AI summaries, translation, speaker ID, and advanced exports.
OpenAI API guide for choosing the right speech architecture for live audio, translation, transcription, speech generation, and audio-capable chat.
Pewbeam is a church presentation app that listens to sermons, detects Bible verse references in real time, and displays the matching passage on screen for smoother live services.
Dictato is a Mac dictation app that transcribes speech to text in any app with an offline, on-device workflow. Includes cleanup, translation, and one-time purchase.
Voicenotes is an AI note-taking and meeting recording app that transcribes conversations, creates summaries and action items, and makes past recordings searchable with Ask AI.