Sprinklr Orchestrated ElevenLabs Speech Engine Configuration
Updated
Voice-enabled AI Agents can be configured using the Enable Voice option while creating or editing an agent. Once voice is enabled, the Voice Stack setting determines which speech engine powers the agent's voice interactions.
The following Voice Stack options are available:
- Sprinklr Stacked Architecture: The standard voice bot infrastructure.
- Sprinklr Orchestrated ElevenLabs Speech Engine: Uses ElevenLabs for both Automatic Speech Recognition (ASR) and Text-to-Speech (TTS).
- Sprinklr Orchestrated Native Speech Engine: Uses Sprinklr’s native speech processing capabilities.
The selected Voice Stack determines the configuration options available during setup, as each engine exposes different speech and voice controls.

Verify Voice AI Access
To confirm that Voice AI functionality is available in your environment:
- Verify that the AI+ Studio → Voice Workflow permission is available under Global Roles or Workspace Roles.
- Assign the permission to the required users or user groups.
- Once assigned, users can access AI+ Studio → AI Use Cases → Sprinklr Service → Voice AI.
If the permission is unavailable, submit an enablement request and provide:
- Environment
- Partner ID
- Language
- Any speech providers that should be excluded
Configure an ElevenLabs Speech Engine Deployment
To create a speech engine deployment using ElevenLabs:
- Open AI+ Studio.
- Select AI Use Cases.
- Navigate to Sprinklr Service → Voice Workflow → Speech Engine.
- Select + Deployment.
- Provide the following information:
Field | Description |
Name | Unique name for the deployment |
Description | Purpose of the deployment |
Orchestration | Select Sprinklr Orchestrated ElevenLabs Speech Engine |
Language | Language used for voice interactions |
GenAI Configuration | Associated GenAI configuration |
Select Next to configure speech settings.
Configure Speech-to-Text, Text-to-Speech, and Speech Engine settings.
Select Save.
After deployment, the speech profile becomes available in the AI Agent’s Voice Settings.
Speech-to-Text Configuration
Configure how user speech is converted into text.
Setting | Description |
ASR Provider | Speech recognition provider used for transcription |
Word Boosting | Improves recognition of specific words or phrases |
Noise Suppression | Removes background noise before transcription |
Noise Suppression Model | Model used for noise filtering. Krisp AI is recommended. |
Noise Suppression Threshold | Controls noise filtering sensitivity. Recommended value: 100 |
Text-to-Speech Configuration
Configure how the AI Agent generates spoken responses.
Setting | Description |
Voice | Voice used for responses |
AI Model | TTS model used for speech generation |
Expressive Model | Controls pronunciation of numbers, symbols, and formatting. Available only with V3 conversational models. |
Text Normalisation | Defines how text is processed before speech generation |
System Prompt | Instructions that guide speech generation behaviour |
Pronunciation Dictionaries | Custom pronunciation rules for words and phrases |
Supported TTS Models
When English is selected:
- eleven_flash_v2
- eleven_v3_conversational
- eleven_multilingual_v2
For other supported languages:
- eleven_flash_v2_5
- eleven_v3_conversational
Audio Tags
Audio tags add speaking instructions to generated speech and help make responses more expressive.
Example:
[excited] Great news! Your request has been approved.
Audio tag descriptions:
- Must clearly describe the desired delivery style.
- Must contain fewer than 200 characters.
- Are supported only with V3 conversational TTS models.
Speech Engine Configuration
These settings control conversational flow and turn-taking.
Setting | Description | Recommended Value |
Turn Model | Determines how speaker turns are detected | v3 |
Eagerness | Controls how quickly the agent responds | Normal |
Take Turn After Silence | Silence duration before the agent responds | 7 seconds |
End Conversation After Silence | Silence duration before ending the conversation | 290 seconds |
Spelling Patience | Wait time while users spell words | Auto |
Ignore Interruption Terms | Prevents specific phrases from triggering interruptions | As required |
Additional Controls
- Enable Filtering of Background Speech: Reduces unintended speech detection.
- Enable Speculative Turn: Allows the agent to prepare responses before the speaker finishes.
- Enable Re-Transcribing Audio on Turn Timeout: Improves transcription accuracy after timeouts.
- Enable Transcription of User Speech on Disabled Interruptions: Continues transcription even when interruptions are disabled.
Sprinklr Orchestrated ElevenLabs Speech Engine
The Sprinklr Orchestrated ElevenLabs Speech Engine integrates ElevenLabs speech services into the voice experience. ElevenLabs handles both:
- Automatic Speech Recognition (ASR)
- Text-to-Speech (TTS)
This integration provides access to advanced speech capabilities, including expressive speech generation, audio tags, pronunciation controls, and enhanced conversational tuning.
Pronunciation Dictionaries
Pronunciation dictionaries allow administrators to define how specific words are spoken.
Common use cases include:
- Product names
- Brand names
- Acronyms
- Industry-specific terminology
- Custom business vocabulary
This provides greater control over speech output without requiring backend changes or development support.
Conversation Tuning Settings
The ElevenLabs Voice Stack provides additional controls for conversational pacing:
- Turn Eagerness controls how readily the agent responds.
- Wait Time After User Silence controls how long the agent waits before replying.
- Interruption Ignore Terms allow up to 50 common acknowledgement phrases, such as "okay" or "got it", to be ignored during interruption detection.
Recommended Settings
Setting | Recommended Value |
Noise Suppression Threshold | 100 |
Turn Eagerness | Normal |
Spelling Patience | Auto |
Turn Model | v3 |
Take Turn After Silence | 7 seconds |
End Conversation After Silence | 290 seconds |
These settings help create more natural and predictable conversational experiences.
Voice Stack vs. Legacy Speech Profiles
Previous voice configurations were managed primarily through speech profiles, which offered basic language and voice selection capabilities.
The new Voice Stack experience provides additional self-service configuration options, including:
- Voice Stack selection
- Voice selection
- TTS model configuration
- Language settings
- Audio tags
- Pronunciation management
- Advanced conversational controls
This expanded configuration experience reduces reliance on engineering teams and gives administrators greater flexibility when designing voice experiences.
Important
The following models are scheduled for deprecation:
- TTS Model: eleven_turbo_v2
- ASR Model: elevenlabs