Sprinklr Speech Engine

Updated 

The Sprinklr Speech Engine is Sprinklr’s native speech processing solution for voice-enabled AI Agents. It is available as a Voice Stack option and enables agents to process spoken customer interactions using configured speech-to-text (STT) and text-to-speech (TTS) capabilities.

When voice is enabled for an AI Agent, you can select one of the following Voice Stack options:

  • Legacy: The original voice bot infrastructure.
  • ElevenLabs Speech Engine: Uses ElevenLabs for speech recognition and speech synthesis.
  • Sprinklr Speech Engine: Uses Sprinklr’s native speech processing capabilities.

Each Voice Stack provides a different set of voice configuration options.

Enable the Sprinklr Speech Engine

To use the Sprinklr Speech Engine:

  1. Create a new AI Agent or open an existing AI Agent.
  2. Enable the Enable Voice option.
  3. From the Voice Stack list, select Sprinklr Speech Engine.

After the Voice Stack is selected, the relevant voice configuration options become available for the agent.

Configure Voice Settings

Voice-enabled AI Agents require a speech profile that defines how speech is processed during conversations.

To configure voice settings:

  1. Navigate to Conversation Settings > Voice Settings.
  2. Select an existing Speech Profile or create a new one.
  3. Save the configuration.

Note: The Voice Settings section is displayed only when voice capabilities are enabled for the AI Agent.

​

About Speech Profiles

A Speech Profile contains the speech configuration used by a voice-enabled AI Agent. It defines the settings used for:

  • Speech-to-Text (STT)
  • Text-to-Speech (TTS)

For the Sprinklr Speech Engine, speech profiles are created and managed in AI Studio and then assigned to the AI Agent.

Once configured, the Speech Profile can be selected under Conversation Settings > Voice Settings.

Create a Speech Profile

If no speech profile exists:

  1. Open Voice Settings.
  2. Select Create Speech Profile.
  3. Configure the required profile settings.

Profile Settings

Field

Description

Profile Name

Name used to identify the speech profile

Speech Configs

Collection of STT and TTS configurations used by the profile

Speech Config

Speech configuration associated with the profile. The configuration marked as Default is used by default.

Language

Language supported by the speech configuration

Configure Text-to-Speech (TTS)

Text-to-Speech settings determine how the AI Agent's responses are converted to audio.

TTS Settings

Setting

Description

TTS Source

Provider used for speech generation

Speech Rate

Speed of generated speech

Volume

Output volume level

Pitch

Voice pitch level

Sample Text

Test text used to preview the configured voice settings

Use sample text to validate the selected voice and speech characteristics before publishing the configuration.