Speech-to-Text (STT) Configuration for Voice AI Agent

Updated 

Speech-to-Text (STT) converts spoken language into text that can be processed by an AI Agent. It enables the agent to understand customer input and determine the appropriate response or action.

STT configuration allows you to control how speech is transcribed, including the language, speech recognition provider, AI model, keyword recognition, multilingual processing, and transcript handling rules. These settings help improve transcription accuracy, especially when conversations include business-specific terminology, product names, codes, or structured information.

Steps to Access STT Configuration

To create or manage an STT configuration:

  1. Open AI+ Studio.
  2. Select AI Use Cases.
  3. Navigate to Sprinklr Service > Voice Workflow > STT.
  4. Select + Configuration.
  5. Configure the required settings and select Save.

Basic Configuration

Configure the foundational settings for the STT profile.

Field

Description

Name

Unique name for the STT configuration

Description

Purpose of the configuration

Language

Language that the STT engine should recognise and transcribe

AI Model Settings

The AI Model Settings section determines how audio is processed and converted into text.

Field

Description

Provider

Speech-to-text provider used for transcription

Model

Speech recognition model available for the selected provider

Note: When Sprinklr In-House is selected as the provider, multiple STT models may be available for selection.


Keyword Boosting

Keyword Boosting improves recognition accuracy for important words and phrases.

This feature is particularly useful for:

  • Product names
  • Brand names
  • Acronyms
  • Technical terminology
  • Frequently used business terms

The Default Keyword List contains keywords that are prioritised when no call-specific keyword list is active.

Use keyword boosting when accurate recognition of specific terms is critical for downstream workflows and AI Agent responses.


Multilingual Speech-to-Text

Multilingual STT allows the AI Agent to recognise and transcribe speech in multiple languages.

Configuration Options

Field

Description

Enable Multilingual Speech-to-Text Conversion

Enables multilingual speech recognition

Language

Additional language that the AI Agent can recognise

Keyword List

Keywords used to identify language-specific speech

This feature is useful for deployments that support customers across multiple languages.


Noise Suppression

Noise Suppression removes unwanted background noise before speech is processed by the STT engine.

This improves transcription quality in environments with background conversations, ambient noise, or audio interference.

Configuration Options

Field

Description

Model

Noise suppression model used to filter audio

Noise Suppression Threshold

Controls the level of background noise filtering

Noise Suppression Threshold

The threshold can be configured between 0 and 100:

  • Lower values retain more background audio.
  • Higher values apply stronger noise filtering.

Adjust the setting based on the quality of the audio source and expected environment.


Transcription Controls

Transcription Controls modify the transcript generated by the STT engine before it is passed to the AI Agent or application.

These controls help standardise transcripts and improve consistency across downstream processes.

The available controls include:

  • Text Normalisation
  • Word Replacements
  • Regex Replacement


Text Normalisation

Text Normalisation standardises recognised content into consistent formats.

Common use cases include:

  • Dates
  • Numbers
  • Email addresses
  • Other supported text patterns

Enable Text Normalisation and select the content types that should be normalised.

Example

A spoken date or number can be automatically converted into a standard format, making it easier for applications to process the transcript consistently.


Word Replacements

Word Replacements allow specific words or phrases to be replaced with predefined alternatives.

This is useful when:

  • Terms are frequently transcribed incorrectly.
  • Multiple variations of a term need to be standardised.
  • Organisation-specific terminology should be used consistently.
  • Preferred values should replace recognised terms.

Configure Word Replacements

  1. Enable Word Replacements.
  2. Download the replacement template.
  3. Add source terms and their replacement values.
  4. Upload the completed template.

The configured mappings are automatically applied during transcript processing.


Regex Replacement

Regex Replacement uses pattern matching to identify and replace structured content within transcripts.

Unlike Word Replacements, which target specific words, Regex Replacement works with values that follow defined patterns.

Common use cases include:

  • Ticket numbers
  • Reference IDs
  • Account numbers
  • Case IDs
  • Product codes

Configuration Options

Field

Description

Pattern

Regular expression used to identify matching text

Replace With

Text that replaces the matched value

You can create multiple regex rules as required.

Example

If ticket IDs follow a standard format, a regex rule can identify and transform those IDs into a consistent format before the transcript is processed further.


How Transcription Controls Work

The transcription flow follows this sequence:

The transcription process can be understood as follows:

Voice Input → STT Provider → Transcribed Text → Transcription Controls → Processed Transcript → AI Agent/Application

The STT provider first converts the user's speech into text. The configured transcription controls can then normalize the text, replace specific terms, or apply pattern-based replacements before the processed transcript is used by the application.

The STT provider first converts speech into text. The configured transcription controls then apply normalisation and replacement rules before the transcript is passed to the AI Agent or application.

Preview Configuration

Preview Configuration allows you to validate STT settings before deploying them in a production voice workflow.

Steps to Test with an Audio Sample

You can test transcription accuracy using recorded audio.

  1. Select a Keyword List for Testing.
  2. Record an audio sample.
  3. Process the audio using the configured STT settings.
  4. Review the generated transcript.

Transcript Preview

The Transcript Preview section displays the generated transcript from a recorded audio sample or selected test case.

Use the preview to verify:

  • Provider selection
  • Model selection
  • Keyword boosting results
  • Multilingual recognition behaviour
  • Text normalisation output
  • Word replacements
  • Regex replacements
  • Overall transcription quality

If no test audio is available, the preview displays No transcripts generated yet.