Text-to-Speech (TTS) Configuration for Voice AI Agent

Updated 

Text-to-Speech (TTS) converts AI Agent responses into natural-sounding speech for voice interactions. TTS settings allow you to configure the voice, language, speech model, pronunciation rules, and voice characteristics used when responding to customers. You can also use the Preview Configuration panel to generate and listen to a sample audio output before saving the configuration.

Steps to Access Text-to-Speech Configuration

  1. Click the New Tab icon.
  2. Under Platform Modules, select AI+ Studio from Learn.
  3. On the AI+ Studio page, click AI Use Cases.
  4. In the Choose your use-case window, select the Sprinklr Service tab.
  5. Under Voice Workflow, navigate to TTS.
  6. On the Text To Speech page, click + Configuration.
  7. In the Text To Speech Configuration window, configure the required settings.
  8. Click Save.

TTS Configuration

Basic Details

Field

Description

Name*

Enter a unique name for the TTS configuration.

Description

Provide information about the purpose or intended use of the configuration.

Voice Model

The Voice Model section determines the voice used to convert AI Agent responses into speech.

  • Select a voice for spoken responses.
  • The selected voice defines characteristics such as identity, tone, and speaking style.
  • Use Explore Voices to review available options and choose the most suitable voice for your use case.

Voice Settings

Configure the following settings to control speech generation:

Field

Description

Language*

Specifies the language used for generated speech.

Model

Selects the speech model used to generate audio output. Available options depend on the chosen voice provider.

Text Normalisation Mode

Defines how numbers, dates, symbols, and other supported formats are interpreted before conversion to speech.

Stability

Controls the consistency of voice output.

Style

Controls the level of expressiveness and stylistic variation.

Speed

Controls how quickly the speech is delivered.

Similarity Boost

Controls how closely the generated speech matches the selected voice characteristics.

Voice Controls

Use the sliders to fine-tune voice output:

  • Stability: Adjusts voice consistency.
  • Style: Adjusts expressiveness and variation.
  • Speed: Adjusts speaking pace.
  • Similarity Boost: Adjusts how closely the generated voice matches the selected voice profile.

The optimal settings depend on your use case and desired customer experience.

Pronunciation Controls

Pronunciation Controls help customise how specific content is spoken. Use these settings when the default pronunciation does not produce the desired output.

Text Normalisation

Enable Text Normalisation to control how symbols, numbers, dates, and other special formats are spoken.

  • Turn on Text Normalisation.
  • Use Apply Rules For to specify the content categories that require normalisation.
  • This ensures that values such as dates, numbers, and symbols are spoken naturally instead of being read character by character.

Word Replacements

Word Replacements allow you to substitute specific words or phrases before generating speech.

Typical use cases include:

  • Correcting pronunciation of specific words.
  • Standardising pronunciation of business terminology.
  • Expanding abbreviations or product names into preferred spoken forms.
  • Ensuring consistent pronunciation across interactions.

Configure Word Replacements

  1. Enable Word Replacements.
  2. Click Download Template.
  3. Add the required word-to-phrase mappings.
  4. Upload the completed template using Upload.

The configured replacements are applied before speech generation.

Regex Replacement

Regex Replacement uses pattern matching to replace text before it is converted into speech.

This is useful for content with predictable formats, such as:

  • Ticket numbers
  • IDs
  • Reference numbers
  • Codes
  • Other structured values

Configure Regex Replacement

Field

Description

Pattern

Enter the regular expression used to identify content for replacement.

Replacement With

Enter the text that replaces the matched pattern.

• Click the + icon to add additional regex replacement rules.

Preview Configuration

The Preview Configuration panel allows you to test TTS settings before using them in live customer interactions.

Text for Preview

Use this section to enter or generate sample text for testing.

You can:

  • Enter custom text.
  • Select one of the suggested scenarios:

    • Transaction Summary
    • Upsell / Offer
    • Deal with Angry Customer
    • Greet the Customer
    • Put the Customer on Hold
    • Repeat and Confirm PII Details
  • Click Generate to create an audio preview.

Preview text can include numbers, dates, email addresses, URLs, and other content types to validate pronunciation and speech behaviour.

Audio Output

The Audio Output section displays the generated preview audio.

  • After configuring TTS settings and entering preview text, click Generate.
  • If no preview has been generated, the section displays No audio to preview yet.

Use the audio preview to verify:

  • Voice selection
  • Speech speed
  • Voice characteristics
  • Pronunciation
  • Text normalisation
  • Word replacements
  • Regex replacements

• •