Voice, text, and visual interaction
The agent combines spoken interaction, written information, and visual content in the same conversation instead of forcing customers to use one medium throughout.
Sierra’s multimodal AI agents combine voice, text, and visual interfaces for customer service, helping teams compare options and make selections in one conversation.
Sierra’s multimodal agents are conversational AI agents that combine voice, text, and visual interfaces in a single customer interaction. Rather than keeping every step in one format, the agent adapts the interface to the task: customers can explain a need by voice, compare options visually, and use text when they need a reference.
The system is designed for customer-service workflows where conversation and structured interaction need to work together. A customer can discuss a disrupted flight, review alternate itineraries, select a seat from a map, or complete a form without restarting the conversation or repeating details. Teams build the agent once, add reusable visual components, and deploy the experience across the channels where the agent is available.
The agent combines spoken interaction, written information, and visual content in the same conversation instead of forcing customers to use one medium throughout.
Sierra agents can select voice for explanations, visuals for side-by-side comparisons, or text for information customers may want to reference later.
Teams can present product cards and comparison tables so customers can review options together rather than relying on a representative to describe them one at a time.
Calendars, seat maps, and forms can be embedded in the conversation, allowing customers to make selections or provide details directly in the interface.
Teams design and host their own components, controlling how they look, what they show, and when they change. Updates are reflected wherever the component is used without separate platform versions.
Larger interfaces can expand to full screen for calendars, long comparison tables, multi-step forms, and other content that needs additional room.
An airline agent can show alternate flights with departure times, layovers, and prices after discussing a disrupted itinerary, then continue after the customer selects an option.
A customer can view a seat map during an airline conversation and tap the preferred seat instead of describing its location verbally.
A telecom or retail agent can discuss requirements by phone and show models, colors, storage sizes, and monthly rates for side-by-side comparison before the customer chooses.
A conversational workflow can use calendars for scheduling and multi-step forms when the customer needs to provide several pieces of information in a larger interface.
Sierra’s multimodal agents combine voice, text, and visual components in one conversation. The agent can choose among these modes based on what the customer needs at each point in the interaction.
A team builds the agent and can deploy it across the channels where the agent lives. Visual components can be reused across those surfaces rather than maintained as separate versions for each platform.
Teams can use Sierra’s MCP UI integration to place interactive product cards, comparison tables, calendars, and forms directly in a conversation. The team designs and hosts these components and controls their appearance, content, and update behavior.
Customers can switch between speaking, reading, viewing options, and tapping a selection without restarting the conversation or repeating previously provided information. Components that need more space can expand to full screen.
Flunkey is a voice-first productivity layer for Windows that turns speech into text, AI outputs, and remembered context without leaving your app.
An All-In-One AI Platform that combines tools for image, video, voice, writing, and chat to enhance creativity and collaboration.
Gemma AI is a phone call reminder app that calls you with scheduled reminders instead of push notifications, with Google Calendar sync and natural voice interaction.
Wallie is an open-source AI streamer that watches your screen, hears chat, and delivers live commentary in a configurable persona. Runs locally with your own keys.
Claude Overlay is a Windows desktop overlay for Claude Code that reads your screen, so you can ask questions, inspect content, and request edits without leaving the app. Uses your existing Claude subscription via Claude CLI.
SpeakoFlow is a free, open-source desktop voice-to-text app for Windows, macOS, and Linux. Dictate into any app, use Flow, cleanup, translation, and a screen-aware assistant.