Configuring Voice AI STT (Speech To Text)
Updated
The Voice AI Speech-to-Text (STT) configuration allows you to control how spoken audio is captured, understood, and converted into text across voice interactions. It defines transcription accuracy, speech detection behavior, vocabulary handling, and post-processing of transcripts.
STT configurations are used in:
- Voice AI Agents (voicebots)
- IVR flows
- Real-Time Streaming (RTS) for agent-customer conversations
This article is intended for contact center administrators and voice solution designers configuring voice intelligence.
For more information on permissions required to view and access TTS Configuration page, refer to the Getting Started with TTS and STT article.
Accessing STT Configuration
Navigate to AI+ Studio from the Sprinklr launchpad.
Click the AI Use Cases card.

Navigate to Sprinklr Service tab.
Under Voice Workflow, select STT. The Speech-to-Text Configurations page appears.

Create a Speech-to-Text Configuration
In the Speech-to-Text Configuration page, click + Configurations to configure the following sections:
The Text-to-Speech configuration page appears.

Basic Details
- Name (required) - Enter a unique name for the STT configuration.
- Description - Provide a brief description for internal reference.
- Language (required) - Select the language for transcription.
This determines the available providers and models.
AI Model Settings
This section defines the provider and model used to transcribe audio.
Provider (required) - Select a Speech-to-Text provider from the list. Depending on your account configuration, available providers may include:
Google Chirp
Sprinklr In-House
Eleven Labs
Azure Speech
Azure OpenAI
Model (provider-dependent)
The selected provider determines the underlying transcription model. Models may differ in:
Accuracy
Supported languages
Latency
Domain specialization
When Provider is set to Sprinklr In-House, you can select from the following available transcription models. The model list displays only supported transcription models registered for your environment.
Sprinklr AI Whisper
Wav2vec2
Whisper
Note: Contact Sprinklr Support to change or add any In-House model.
System Prompt (if supported by the provider)
Provide additional context to improve transcription quality.
For example:
Industry-specific terminology
Expected phrases or formats
Context about the conversation domain
Tip: Use system prompts when working with specialized domains such as banking, healthcare, or technical support.
Keyword Lists (Keyword Boosting)
Use Keyword Boosting to improve recognition of domain-specific terms.
- Default Keyword List - Select a predefined keyword list.
Keyword lists help the system correctly recognize:
- Brand names
- Product names
- Industry-specific vocabulary
- Case linked variables
- Profile linked variables
- User level keywords (example, name of customer)
Example: Adding product names or company-specific terms improves transcription accuracy in customer conversations.
Multi Lingual Speech-To-Text
The Enable Multilingual Speech-to-Text Conversion option allows a single Speech-to-Text deployment to process conversations that may contain multiple languages.
When enabled, you can:
- Select a primary language for transcription.
- Add one or more additional languages supported by the selected provider and model.
- Associate a keyword boosting list with each language.
- Configure language-specific settings, such as accents, where supported by the selected provider.
The system dynamically displays available languages and supported accents based on the selected provider and model. Invalid language or accent combinations cannot be saved.
Example
A contact center serving customers in India can create a single STT deployment with:
- Primary language: English
- Additional language: Hindi
- English keyword boosting list
- Hindi keyword boosting list
During a conversation, based on the language passed, the configuration switches languages and applies the corresponding keyword boosting configuration to improve transcription quality.
Note: The number of supported languages depends on the capabilities and restrictions of the selected STT provider and model. Available languages, accents, and limits are populated dynamically during configuration.
Noise Suppression
Enable Noise Suppression to reduce background noise in incoming audio before transcription.
Noise Suppression can help improve speech recognition accuracy in environments where callers may experience:
- Office background noise
- Public environments
- Call center ambient noise
- Multiple speakers in close proximity
- General environmental disturbances
Configure Noise Suppression
Field | Description |
Model | Select the noise suppression model used to filter incoming audio. Currently, the In House Sprinklr model is supported. |
Noise Suppression Threshold | Determines the level of noise reduction applied to incoming audio. Higher values provide stronger noise filtering, while lower values preserve more original audio. |
Transcription Controls
This section defines how the final transcript is formatted before it is consumed by downstream systems.
Text Normalization
Enable Text Normalization to standardize transcription output.
This ensures consistent formatting of:
- Dates
- Numbers
- Email addresses
- Structured phrases
Example:
“twenty fifth May twenty twenty six” → “05/25/2026”
Word Replacements
Enable Word Replacements to substitute commonly mistranscribed words with preferred alternatives.
Use this to:
- Correct recurring transcription errors
- Standardize terminology
Example:
“sprinkler” → “Sprinklr”
Regex Replacement
Enable Regex Replacement to apply pattern-based transformations to structured text.
Use this for:
- Ticket numbers
- Account IDs
- Reference codes
These rules are applied after transcription to improve data consistency.
Preview Configuration
Use the Preview Configuration panel (right side of the screen) to test the STT setup.
Audio Sample
- Record or upload an audio sample
- Optionally select Additional Hints for Testing to guide the system
Transcript Preview
- View the generated transcript after processing
- Validate accuracy and formatting
Note: If no sample is recorded, the transcript preview remains empty.
Where This Configuration Is Used
Once created, the STT configuration is available in Speech Profiles under:
Voice Care → Voice Settings → Speech Profiles → STT Source
Important Notes
- STT configurations are reusable across deployments
- Changes to a configuration affect all linked speech profiles
- Available providers and features depend on language selection
Availability Note
STT Configuration is available only for new deployments created after this feature release. Existing deployments are not automatically migrated.
Related Articles