Sprinklr Orchestrated ElevenLabs Speech Engine Configuration

Updated 

Voice-enabled AI Agents can be configured using the Enable Voice option while creating or editing an agent. Once voice is enabled, the Voice Stack setting determines which speech engine powers the agent's voice interactions.

The following Voice Stack options are available:

  • Sprinklr Stacked Architecture: The standard voice bot infrastructure.
  • Sprinklr Orchestrated ElevenLabs Speech Engine: Uses ElevenLabs for both Automatic Speech Recognition (ASR) and Text-to-Speech (TTS).
  • Sprinklr Orchestrated Native Speech Engine: Uses Sprinklr’s native speech processing capabilities.

The selected Voice Stack determines the configuration options available during setup, as each engine exposes different speech and voice controls.

Verify Voice AI Access

To confirm that Voice AI functionality is available in your environment:

  1. Verify that the AI+ Studio → Voice Workflow permission is available under Global Roles or Workspace Roles.
  2. Assign the permission to the required users or user groups.
  3. Once assigned, users can access AI+ Studio → AI Use Cases → Sprinklr Service → Voice AI.
  4. If the permission is unavailable, submit an enablement request and provide:

    • Environment
    • Partner ID
    • Language
    • Any speech providers that should be excluded

Configure an ElevenLabs Speech Engine Deployment

To create a speech engine deployment using ElevenLabs:

  1. Open AI+ Studio.
  2. Select AI Use Cases.
  3. Navigate to Sprinklr Service → Voice Workflow → Speech Engine.
  4. Select + Deployment.
  5. Provide the following information:

Field

Description

Name

Unique name for the deployment

Description

Purpose of the deployment

Orchestration

Select Sprinklr Orchestrated ElevenLabs Speech Engine

Language

Language used for voice interactions

GenAI Configuration

Associated GenAI configuration

  1. Select Next to configure speech settings.

  2. Configure Speech-to-Text, Text-to-Speech, and Speech Engine settings.

  3. Select Save.

After deployment, the speech profile becomes available in the AI Agent’s Voice Settings.

Speech-to-Text Configuration

Configure how user speech is converted into text.

Setting

Description

ASR Provider

Speech recognition provider used for transcription

Word Boosting

Improves recognition of specific words or phrases

Noise Suppression

Removes background noise before transcription

Noise Suppression Model

Model used for noise filtering. Krisp AI is recommended.

Noise Suppression Threshold

Controls noise filtering sensitivity. Recommended value: 100

Text-to-Speech Configuration

Configure how the AI Agent generates spoken responses.

Setting

Description

Voice

Voice used for responses

AI Model

TTS model used for speech generation

Expressive Model

Controls pronunciation of numbers, symbols, and formatting. Available only with V3 conversational models.

Text Normalisation

Defines how text is processed before speech generation

System Prompt

Instructions that guide speech generation behaviour

Pronunciation Dictionaries

Custom pronunciation rules for words and phrases

Supported TTS Models

When English is selected:

  • eleven_flash_v2
  • eleven_v3_conversational
  • eleven_multilingual_v2

For other supported languages:

  • eleven_flash_v2_5
  • eleven_v3_conversational

Audio Tags

Audio tags add speaking instructions to generated speech and help make responses more expressive.

Example:

[excited] Great news! Your request has been approved.

Audio tag descriptions:

  • Must clearly describe the desired delivery style.
  • Must contain fewer than 200 characters.
  • Are supported only with V3 conversational TTS models.

Speech Engine Configuration

These settings control conversational flow and turn-taking.

Setting

Description

Recommended Value

Turn Model

Determines how speaker turns are detected

v3

Eagerness

Controls how quickly the agent responds

Normal

Take Turn After Silence

Silence duration before the agent responds

7 seconds

End Conversation After Silence

Silence duration before ending the conversation

290 seconds

Spelling Patience

Wait time while users spell words

Auto

Ignore Interruption Terms

Prevents specific phrases from triggering interruptions

As required

Additional Controls

  • Enable Filtering of Background Speech: Reduces unintended speech detection.
  • Enable Speculative Turn: Allows the agent to prepare responses before the speaker finishes.
  • Enable Re-Transcribing Audio on Turn Timeout: Improves transcription accuracy after timeouts.
  • Enable Transcription of User Speech on Disabled Interruptions: Continues transcription even when interruptions are disabled.

Sprinklr Orchestrated ElevenLabs Speech Engine

The Sprinklr Orchestrated ElevenLabs Speech Engine integrates ElevenLabs speech services into the voice experience. ElevenLabs handles both:

  • Automatic Speech Recognition (ASR)
  • Text-to-Speech (TTS)

This integration provides access to advanced speech capabilities, including expressive speech generation, audio tags, pronunciation controls, and enhanced conversational tuning.

Pronunciation Dictionaries

Pronunciation dictionaries allow administrators to define how specific words are spoken.

Common use cases include:

  • Product names
  • Brand names
  • Acronyms
  • Industry-specific terminology
  • Custom business vocabulary

This provides greater control over speech output without requiring backend changes or development support.

Conversation Tuning Settings

The ElevenLabs Voice Stack provides additional controls for conversational pacing:

  • Turn Eagerness controls how readily the agent responds.
  • Wait Time After User Silence controls how long the agent waits before replying.
  • Interruption Ignore Terms allow up to 50 common acknowledgement phrases, such as "okay" or "got it", to be ignored during interruption detection.

Recommended Settings

Setting

Recommended Value

Noise Suppression Threshold

100

Turn Eagerness

Normal

Spelling Patience

Auto

Turn Model

v3

Take Turn After Silence

7 seconds

End Conversation After Silence

290 seconds

These settings help create more natural and predictable conversational experiences.

Voice Stack vs. Legacy Speech Profiles

Previous voice configurations were managed primarily through speech profiles, which offered basic language and voice selection capabilities.

The new Voice Stack experience provides additional self-service configuration options, including:

  • Voice Stack selection
  • Voice selection
  • TTS model configuration
  • Language settings
  • Audio tags
  • Pronunciation management
  • Advanced conversational controls

This expanded configuration experience reduces reliance on engineering teams and gives administrators greater flexibility when designing voice experiences.

Important

The following models are scheduled for deprecation:

  • TTS Model: eleven_turbo_v2
  • ASR Model: elevenlabs