Voice, text, and visual interaction
The agent combines spoken interaction, written information, and visual content in the same conversation instead of forcing customers to use one medium throughout.
Sierra’s multimodal agents are conversational AI agents that combine voice, text, and visual interfaces in a single customer interaction. Rather than keeping every step in one format, the agent adapts the interface to the task: customers can explain a need by voice, compare options visually, and use text when they need a reference.
The system is designed for customer-service workflows where conversation and structured interaction need to work together. A customer can discuss a disrupted flight, review alternate itineraries, select a seat from a map, or complete a form without restarting the conversation or repeating details. Teams build the agent once, add reusable visual components, and deploy the experience across the channels where the agent is available.
The agent combines spoken interaction, written information, and visual content in the same conversation instead of forcing customers to use one medium throughout.
Sierra agents can select voice for explanations, visuals for side-by-side comparisons, or text for information customers may want to reference later.
Teams can present product cards and comparison tables so customers can review options together rather than relying on a representative to describe them one at a time.
Calendars, seat maps, and forms can be embedded in the conversation, allowing customers to make selections or provide details directly in the interface.
Teams design and host their own components, controlling how they look, what they show, and when they change. Updates are reflected wherever the component is used without separate platform versions.
Larger interfaces can expand to full screen for calendars, long comparison tables, multi-step forms, and other content that needs additional room.
An airline agent can show alternate flights with departure times, layovers, and prices after discussing a disrupted itinerary, then continue after the customer selects an option.
A customer can view a seat map during an airline conversation and tap the preferred seat instead of describing its location verbally.
A telecom or retail agent can discuss requirements by phone and show models, colors, storage sizes, and monthly rates for side-by-side comparison before the customer chooses.
A conversational workflow can use calendars for scheduling and multi-step forms when the customer needs to provide several pieces of information in a larger interface.
Sierra’s multimodal agents combine voice, text, and visual components in one conversation. The agent can choose among these modes based on what the customer needs at each point in the interaction.
A team builds the agent and can deploy it across the channels where the agent lives. Visual components can be reused across those surfaces rather than maintained as separate versions for each platform.
Teams can use Sierra’s MCP UI integration to place interactive product cards, comparison tables, calendars, and forms directly in a conversation. The team designs and hosts these components and controls their appearance, content, and update behavior.
Customers can switch between speaking, reading, viewing options, and tapping a selection without restarting the conversation or repeating previously provided information. Components that need more space can expand to full screen.
Flunkey 是 Windows 的語音優先生產力工具,將口說內容轉成文字、AI 輔助輸出與可記住的上下文,讓你在不離開當前應用程式的情況下快速記錄想法與待辦。
一個集成圖像、視頻、語音、寫作和聊天工具的全能AI平台,以增強創造力和協作。
Gemma AI 是一款電話提醒 app,會依排程直接致電提醒你,不靠推播通知。支援 Google Calendar 同步與自然對話互動,讓你更直接掌握行程。
Wallie 是開源 AI streamer,可觀看你的螢幕、聆聽聊天室,並以可設定的人設即時生成直播評論;支援本機執行與自有金鑰,適合無真人出鏡、自治直播與即時互動。
Claude Overlay 是適用於 Claude Code 的 Windows 桌面浮動疊加層,可讀取螢幕內容,讓你在不離開目前應用程式下提問、檢視內容並請求編輯。透過既有 Claude 訂閱與 Claude CLI 運作,無需額外 API 金鑰。
SpeakoFlow 是適用於 Windows、macOS 與 Linux 的免費開源桌面語音轉文字應用,支援對任何 App 進行聽寫、Flow 語音寫作、AI 清理、翻譯與螢幕感知助理。