Speech-to-Text (STT) Configuration for Voice AI Agent
Updated
Speech-to-Text (STT) converts spoken language into text that can be processed by an AI Agent. It enables the agent to understand customer input and determine the appropriate response or action.
STT configuration allows you to control how speech is transcribed, including the language, speech recognition provider, AI model, keyword recognition, multilingual processing, and transcript handling rules. These settings help improve transcription accuracy, especially when conversations include business-specific terminology, product names, codes, or structured information.
Steps to Access STT Configuration
To create or manage an STT configuration:
- Open AI+ Studio.
- Select AI Use Cases.
- Navigate to Sprinklr Service > Voice Workflow > STT.
- Select + Configuration.
- Configure the required settings and select Save.
Basic Configuration
Configure the foundational settings for the STT profile.
Field | Description |
Name | Unique name for the STT configuration |
Description | Purpose of the configuration |
Language | Language that the STT engine should recognise and transcribe |
AI Model Settings
The AI Model Settings section determines how audio is processed and converted into text.
Field | Description |
Provider | Speech-to-text provider used for transcription |
Model | Speech recognition model available for the selected provider |
Note: When Sprinklr In-House is selected as the provider, multiple STT models may be available for selection.
Keyword Boosting
Keyword Boosting improves recognition accuracy for important words and phrases.
This feature is particularly useful for:
- Product names
- Brand names
- Acronyms
- Technical terminology
- Frequently used business terms
The Default Keyword List contains keywords that are prioritised when no call-specific keyword list is active.
Use keyword boosting when accurate recognition of specific terms is critical for downstream workflows and AI Agent responses.
Multilingual Speech-to-Text
Multilingual STT allows the AI Agent to recognise and transcribe speech in multiple languages.
Configuration Options
Field | Description |
Enable Multilingual Speech-to-Text Conversion | Enables multilingual speech recognition |
Language | Additional language that the AI Agent can recognise |
Keyword List | Keywords used to identify language-specific speech |
This feature is useful for deployments that support customers across multiple languages.
Noise Suppression
Noise Suppression removes unwanted background noise before speech is processed by the STT engine.
This improves transcription quality in environments with background conversations, ambient noise, or audio interference.
Configuration Options
Field | Description |
Model | Noise suppression model used to filter audio |
Noise Suppression Threshold | Controls the level of background noise filtering |
Noise Suppression Threshold
The threshold can be configured between 0 and 100:
- Lower values retain more background audio.
- Higher values apply stronger noise filtering.
Adjust the setting based on the quality of the audio source and expected environment.
Transcription Controls
Transcription Controls modify the transcript generated by the STT engine before it is passed to the AI Agent or application.
These controls help standardise transcripts and improve consistency across downstream processes.
The available controls include:
- Text Normalisation
- Word Replacements
- Regex Replacement
Text Normalisation
Text Normalisation standardises recognised content into consistent formats.
Common use cases include:
- Dates
- Numbers
- Email addresses
- Other supported text patterns
Enable Text Normalisation and select the content types that should be normalised.
Example
A spoken date or number can be automatically converted into a standard format, making it easier for applications to process the transcript consistently.
Word Replacements
Word Replacements allow specific words or phrases to be replaced with predefined alternatives.
This is useful when:
- Terms are frequently transcribed incorrectly.
- Multiple variations of a term need to be standardised.
- Organisation-specific terminology should be used consistently.
- Preferred values should replace recognised terms.
Configure Word Replacements
- Enable Word Replacements.
- Download the replacement template.
- Add source terms and their replacement values.
- Upload the completed template.
The configured mappings are automatically applied during transcript processing.
Regex Replacement
Regex Replacement uses pattern matching to identify and replace structured content within transcripts.
Unlike Word Replacements, which target specific words, Regex Replacement works with values that follow defined patterns.
Common use cases include:
- Ticket numbers
- Reference IDs
- Account numbers
- Case IDs
- Product codes
Configuration Options
Field | Description |
Pattern | Regular expression used to identify matching text |
Replace With | Text that replaces the matched value |
You can create multiple regex rules as required.
Example
If ticket IDs follow a standard format, a regex rule can identify and transform those IDs into a consistent format before the transcript is processed further.
How Transcription Controls Work
The transcription flow follows this sequence:
The transcription process can be understood as follows:
Voice Input → STT Provider → Transcribed Text → Transcription Controls → Processed Transcript → AI Agent/Application
The STT provider first converts the user's speech into text. The configured transcription controls can then normalize the text, replace specific terms, or apply pattern-based replacements before the processed transcript is used by the application.
The STT provider first converts speech into text. The configured transcription controls then apply normalisation and replacement rules before the transcript is passed to the AI Agent or application.
Preview Configuration
Preview Configuration allows you to validate STT settings before deploying them in a production voice workflow.
Steps to Test with an Audio Sample
You can test transcription accuracy using recorded audio.
- Select a Keyword List for Testing.
- Record an audio sample.
- Process the audio using the configured STT settings.
- Review the generated transcript.
Transcript Preview
The Transcript Preview section displays the generated transcript from a recorded audio sample or selected test case.
Use the preview to verify:
- Provider selection
- Model selection
- Keyword boosting results
- Multilingual recognition behaviour
- Text normalisation output
- Word replacements
- Regex replacements
- Overall transcription quality
If no test audio is available, the preview displays No transcripts generated yet.