Sprinklr Speech Engine
Updated
The Sprinklr Speech Engine is Sprinklr’s native speech processing solution for voice-enabled AI Agents. It is available as a Voice Stack option and enables agents to process spoken customer interactions using configured speech-to-text (STT) and text-to-speech (TTS) capabilities.
When voice is enabled for an AI Agent, you can select one of the following Voice Stack options:
- Legacy: The original voice bot infrastructure.
- ElevenLabs Speech Engine: Uses ElevenLabs for speech recognition and speech synthesis.
- Sprinklr Speech Engine: Uses Sprinklr’s native speech processing capabilities.
Each Voice Stack provides a different set of voice configuration options.
Enable the Sprinklr Speech Engine
To use the Sprinklr Speech Engine:
- Create a new AI Agent or open an existing AI Agent.
- Enable the Enable Voice option.
- From the Voice Stack list, select Sprinklr Speech Engine.
After the Voice Stack is selected, the relevant voice configuration options become available for the agent.
Configure Voice Settings
Voice-enabled AI Agents require a speech profile that defines how speech is processed during conversations.
To configure voice settings:
- Navigate to Conversation Settings > Voice Settings.
- Select an existing Speech Profile or create a new one.
- Save the configuration.
Note: The Voice Settings section is displayed only when voice capabilities are enabled for the AI Agent.
About Speech Profiles
A Speech Profile contains the speech configuration used by a voice-enabled AI Agent. It defines the settings used for:
- Speech-to-Text (STT)
- Text-to-Speech (TTS)
For the Sprinklr Speech Engine, speech profiles are created and managed in AI Studio and then assigned to the AI Agent.
Once configured, the Speech Profile can be selected under Conversation Settings > Voice Settings.
Create a Speech Profile
If no speech profile exists:
- Open Voice Settings.
- Select Create Speech Profile.
- Configure the required profile settings.
Profile Settings
Field | Description |
Profile Name | Name used to identify the speech profile |
Speech Configs | Collection of STT and TTS configurations used by the profile |
Speech Config | Speech configuration associated with the profile. The configuration marked as Default is used by default. |
Language | Language supported by the speech configuration |
Configure Text-to-Speech (TTS)
Text-to-Speech settings determine how the AI Agent's responses are converted to audio.
TTS Settings
Setting | Description |
TTS Source | Provider used for speech generation |
Speech Rate | Speed of generated speech |
Volume | Output volume level |
Pitch | Voice pitch level |
Sample Text | Test text used to preview the configured voice settings |
Use sample text to validate the selected voice and speech characteristics before publishing the configuration.