Universal-3.5 Pro icon

Universal-3.5 Pro

Universal-3.5 Pro is AssemblyAI’s async speech-to-text model for recorded audio and video. It adds native code-switching, speaker diarization, and contextual prompting, with a base price of $0.21 per hour of audio submitted.

Universal-3.5 Pro

Overview

Universal-3.5 Pro is AssemblyAI’s flagship asynchronous speech-to-text model for recorded audio and video. It is part of the Pre-recorded Speech-to-Text API and is positioned as the highest-accuracy model for supported languages and difficult real-world audio.

The launch article highlights three main capabilities: native code-switching across 18 languages, improved speaker diarization, and contextual prompting. The pricing page lists the base rate at $0.21 per hour of audio submitted, with optional add-ons for capabilities such as diarization, medical mode, keyterms prompting, and general prompting.

In practice, the model is aimed at transcription workloads where speech is messy, multilingual, or speaker-heavy. That includes meetings, contact center recordings, captions, searchable archives, and other recorded workflows where accuracy and speaker attribution matter more than live latency.

Core capabilities

Native code-switching across 18 languages

The launch article says the model handles code-switched speech natively across 18 languages, so transcripts can follow the language actually spoken without a separate language pass.

Improved speaker diarization for overlapping speech

AssemblyAI describes this release as its most accurate speaker diarization yet, with the model designed to keep speaker attribution aligned with the transcript even in crosstalk and noisy conversations.

Contextual prompting

Contextual prompting lets you provide domain knowledge or prior context to influence the transcription output, which can improve accuracy on specialized terms or ambiguous audio.

Keyterms prompting for custom vocabulary

The pricing reference lists keyterms prompting for Universal-3.5 Pro with support for up to 1,000 terms, giving teams a way to bias recognition toward project-specific vocabulary.

Pre-recorded transcription model

Universal-3.5 Pro is the async pre-recorded model in AssemblyAI’s speech-to-text lineup and is priced at $0.21 per hour of audio submitted.

Add-on support for specialized workflows

The pricing page shows optional add-ons for this model, including medical mode and speaker diarization, which can be combined with the base transcription rate.

Common use cases

  • Multilingual recorded conversations

    Use Universal-3.5 Pro when a recording contains multiple languages in the same conversation and you want the transcript to preserve the language spoken at each point.

  • Meetings with overlapping speakers

    Use the model for meeting notes or internal discussions where short turns, interruptions, and overlapping voices make speaker attribution difficult.

  • Contact center recordings

    Use it for contact center calls or other customer interactions where accurate speaker labeling affects downstream analysis and QA.

  • Domain-specific transcription

    Use contextual prompting or keyterms prompting when transcribing specialized discussions that include project names, domain terms, or other vocabulary that may be misheard.

  • Batch transcription of recorded media

    Use the async API for archives, captions, or dataset preparation when the audio is already recorded and you can trade latency for transcription quality.

Pros and Cons

Pros

  • Supports native code-switching across 18 languages, which is useful when conversations move between languages mid-sentence or mid-call.
  • Adds contextual prompting and keyterms prompting to help guide transcription toward domain-specific vocabulary.
  • Improves speaker diarization for conversations with crosstalk, interruptions, and short turns.
  • Priced separately from real-time models, with a clear async rate in the pricing reference.
  • Fits a broad set of recorded-audio workflows, including meetings, contact center audio, media, and archives.

Cons

  • It is an async model, so it is not the right fit for live transcription use cases that need streaming output.
  • Several advanced capabilities are billed as add-ons rather than being included in the base transcription rate.
  • The source material does not provide a full list of supported languages, output formats, or setup details for the model itself.

FAQ

What kind of transcription workflow is Universal-3.5 Pro for?

Universal-3.5 Pro is AssemblyAI’s async pre-recorded speech-to-text model. It is designed for recorded audio and video rather than live streaming transcription.

What language support does it offer?

The launch article says it supports native code-switching across 18 languages. The product pages identify it as the highest-accuracy pre-recorded model for supported languages.

Can I guide the model with extra context?

The launch article highlights contextual prompting, and the pricing reference shows general prompting is available as a paid async add-on for Universal-3.5 Pro.

Does it support speaker diarization?

The pricing page lists async speaker diarization for Universal-3.5 Pro as a paid add-on, with standard and experimental variants. The launch article positions diarization as one of the model’s core improvements.

How is Universal-3.5 Pro priced?

The pricing page lists Universal-3.5 Pro at $0.21 per hour for pre-recorded audio submitted. Add-ons such as diarization, medical mode, keyterms prompting, and general prompting are priced separately.

Quick Facts

Category
Pre-recorded speech-to-text API
Platform
AssemblyAI Cloud and self-hosted Voice AI infrastructure
Primary users
Developers building transcription workflows for recorded audio and video
Base price
$0.21/hr for pre-recorded audio submitted
Model API value
universal-3-pro
Source domain
assemblyai.com

Universal-3.5 Pro Alternativen

QuickQuill icon

QuickQuill

QuickQuill is a macOS dictation and transcription app that runs locally on the device. It helps users record meetings, transcribe audio, generate summaries, and export notes without using a cloud service.

Speech to Text Converter icon

Speech to Text Converter

Speech to Text Converter is a browser-based transcription tool for live dictation and uploaded audio or video files. It offers a free tier for short tasks and a Pro plan for unlimited transcription, AI summaries, translation, speaker identification, and advanced exports.

Realtime and audio icon

Realtime and audio

An OpenAI API guide for choosing the right speech architecture for live audio, translation, transcription, speech generation, and audio-capable chat. It helps developers map each speech application to the appropriate session type, endpoint, and connection method.

Pewbeam icon

Pewbeam

Pewbeam is a church presentation app that listens to sermons, detects Bible verse references in real time, and displays the matching passage on screen. It is built for pastors, projection teams, and church media volunteers who want to reduce manual slide control during live services.

Dictato icon

Dictato

Dictato is a Mac dictation app that transcribes speech into text in any app using an on-device, offline workflow. It supports multiple transcription engines, optional cleanup and translation, and a one-time purchase license.

Voicenotes icon

Voicenotes

Voicenotes is an AI note-taking and meeting recording app that transcribes conversations, generates summaries and action items, and makes past recordings searchable with Ask AI. It also supports voice dictation, imported audio on Pro, and multilingual transcription.