Fylom is an AI podcast generator that turns a topic or question into a researched narrated episode. It supports typed or spoken prompts, voice selection, and recurring show-style listening.
Seed Audio AI is a browser-based text-to-speech and voice cloning tool for turning scripts into reviewable voice audio. It is aimed at creators and teams that need drafts for voiceovers, narration, podcasts, lessons, and ads.
Vois is a local AI voice generator studio for Mac and Windows that handles speech generation, editing, mastering, and export on your own machine. It includes voice cloning, platform export presets, and paid plans with unlimited generation.
SKI is a desktop app for Mac and Windows that lets developers talk to AI coding agents and hear responses aloud. It keeps the main voice loop local and offline, with optional meeting transcription and agent-assisted call features.
Chariot is an AI text-to-speech product for generating English, Hindi, and Hinglish speech from text, with low-latency streaming and API access for developers. It is aimed at voice agents, app audio, and other production workflows that need natural-sounding speech.
gstack adds Garry Tan’s open-source specialist personas to live meetings as voice bots. It supports Google Meet, with support also mentioned for Zoom and Teams, and runs through a bring-your-own-brain workflow using a local coding-agent session.
PodcastorAI is an AI podcast studio that turns content into podcast scripts, audio episodes, and video podcasts. It helps creators produce publish-ready shows from topics, documents, URLs, notes, or recordings without a traditional studio setup.
VocalVia is a document-to-podcast tool that converts PDFs, articles, notes, Markdown, and web sources into editable podcast drafts and audio. It is aimed at people who want to review the outline and script before generating the final file.
SpeechifyAI is a voice AI platform for generating speech, cloning voices, and building voice agents. It serves developers who need text-to-speech, multilingual audio, and calling workflows through a single API.
Alvoff Inference is an OpenAI-compatible API for speech-to-text, text-to-speech, embeddings, and chat/code generation. It is built for developers who want to swap in a different base URL, use familiar SDKs, and pay per request.
speech-core is a C++17 library for on-device speech orchestration, including VAD, streaming and batch STT, diarization, TTS, and a voice-agent pipeline. It runs locally and uses optional ONNX Runtime or LiteRT backends for model inference.
Voiser AI Voiceover turns text into spoken audio for voiceovers, with multilingual voice options and style controls for different narration needs. It supports a web studio workflow and shows free, paid, and enterprise paths on the site.
Tico é um assistente de IA para Windows que acompanha o cursor, entende o que está na tela e orienta o usuário por voz. A página indica uso gratuito com limite diário e planos pagos com mais usos e suporte prioritário.
Yeta AI is a browser-based tool that translates and dubs public YouTube videos in real time using AI voices. It is designed for watching tutorials, lectures, and other long-form videos in more than 10 languages without relying on subtitles.
Morph is a web-based reading platform for public-domain classics that combines text, synced narration, and an AI assistant. It helps readers switch between reading and listening, browse a curated library, and get book-specific help without leaving the page.
FlowSpeech is a context-aware text-to-speech studio that turns scripts and uploaded files into human-like audio. It offers multiple generation modes, pause and emotion control, and a free plan alongside paid tiers.
xAI’s Grok Speech to Text and Text to Speech APIs let developers add transcription and speech generation to apps through REST and WebSocket endpoints. The product supports multilingual STT, expressive TTS, and usage-based pricing.
Gemini 3.1 Flash TTS is Google’s preview text-to-speech model for generating expressive AI speech with fine-grained control over style and delivery. It is available across the Gemini API, Google AI Studio, Vertex AI, and Google Vids.
Guardrails 2.0 is ElevenLabs’ control layer for ElevenAgents, designed to keep AI voice agents on-topic, policy-aligned, and safer to deploy in production. It is built for teams using voice agents in support, sales, marketing, reception, and internal workflows.
Official HeyGen API documentation for building AI avatar videos, translations, lipsync, and interactive video-agent sessions. It supports direct API use plus MCP and CLI-style workflows for developers and AI agents.