Sierra icon

Sierra

Sierra’s multimodal agents combine voice, text, and visual interfaces in a single customer-service conversation. They help teams build interactions where customers can discuss needs, compare options, and make selections without restarting or repeating the conversation.

Sierra

Multimodal conversations for customer service

Sierra’s multimodal agents are conversational AI agents that combine voice, text, and visual interfaces in a single customer interaction. Rather than keeping every step in one format, the agent adapts the interface to the task: customers can explain a need by voice, compare options visually, and use text when they need a reference.

The system is designed for customer-service workflows where conversation and structured interaction need to work together. A customer can discuss a disrupted flight, review alternate itineraries, select a seat from a map, or complete a form without restarting the conversation or repeating details. Teams build the agent once, add reusable visual components, and deploy the experience across the channels where the agent is available.

Core features

Voice, text, and visual interaction

The agent combines spoken interaction, written information, and visual content in the same conversation instead of forcing customers to use one medium throughout.

Context-aware mode selection

Sierra agents can select voice for explanations, visuals for side-by-side comparisons, or text for information customers may want to reference later.

Interactive product comparisons

Teams can present product cards and comparison tables so customers can review options together rather than relying on a representative to describe them one at a time.

Embedded interactive components

Calendars, seat maps, and forms can be embedded in the conversation, allowing customers to make selections or provide details directly in the interface.

Centralized component management

Teams design and host their own components, controlling how they look, what they show, and when they change. Updates are reflected wherever the component is used without separate platform versions.

Full-screen expansion

Larger interfaces can expand to full screen for calendars, long comparison tables, multi-step forms, and other content that needs additional room.

Practical use cases

  • Flight disruption and rebooking

    An airline agent can show alternate flights with departure times, layovers, and prices after discussing a disrupted itinerary, then continue after the customer selects an option.

  • Seat selection

    A customer can view a seat map during an airline conversation and tap the preferred seat instead of describing its location verbally.

  • Plan or product selection

    A telecom or retail agent can discuss requirements by phone and show models, colors, storage sizes, and monthly rates for side-by-side comparison before the customer chooses.

  • Scheduling and form completion

    A conversational workflow can use calendars for scheduling and multi-step forms when the customer needs to provide several pieces of information in a larger interface.

Pros and Cons

Pros

  • Combines voice, text, and visuals within one continuous conversation.
  • Lets customers compare options and make selections through interactive components rather than text or spoken descriptions alone.
  • Supports reusable components across the agent’s deployed surfaces, reducing the need to maintain separate versions.
  • Allows teams to control the design, content, and update behavior of embedded components.
  • Can expand larger interfaces such as calendars and multi-step forms to full screen.

Cons

  • The source does not provide pricing, plan limits, or detailed channel availability.
  • Teams are responsible for designing and hosting the interactive components used in the experience.
  • The source describes the concept and workflows but does not provide implementation-level technical specifications.

FAQ

What are Sierra’s multimodal agents?

Sierra’s multimodal agents combine voice, text, and visual components in one conversation. The agent can choose among these modes based on what the customer needs at each point in the interaction.

How are multimodal agents deployed?

A team builds the agent and can deploy it across the channels where the agent lives. Visual components can be reused across those surfaces rather than maintained as separate versions for each platform.

What interactive components can be added to a conversation?

Teams can use Sierra’s MCP UI integration to place interactive product cards, comparison tables, calendars, and forms directly in a conversation. The team designs and hosts these components and controls their appearance, content, and update behavior.

How does the experience handle changes between voice, text, and visuals?

Customers can switch between speaking, reading, viewing options, and tapping a selection without restarting the conversation or repeating previously provided information. Components that need more space can expand to full screen.

Quick Facts

Product
Sierra multimodal agents
Category
Conversational AI and customer service
Interaction modes
Voice, text, and visuals
Interactive UI
Product cards, comparison tables, calendars, forms, and seat maps
Deployment model
Build once and deploy across the agent’s available channels
Provider
Sierra