Configuring Voice AI STT (Speech To Text)

Updated 

The Voice AI Speech-to-Text (STT) configuration allows you to control how spoken audio is captured, understood, and converted into text across voice interactions. It defines transcription accuracy, speech detection behavior, vocabulary handling, and post-processing of transcripts.

STT configurations are used in:

  • Voice AI Agents (voicebots)
  • IVR flows
  • Real-Time Streaming (RTS) for agent-customer conversations

This article is intended for contact center administrators and voice solution designers configuring voice intelligence.

​

For more information on permissions required to view and access TTS Configuration page, refer to the Getting Started with TTS and STT article.

Accessing STT Configuration

  1. Navigate to AI+ Studio from the Sprinklr launchpad.

  2. Click the AI Use Cases card.

    ​

  3. Navigate to Sprinklr Service tab.

  4. Under Voice Workflow, select STT. The Speech-to-Text Configurations page appears.

    ​

    ​

Create a Speech-to-Text Configuration

In the Speech-to-Text Configuration page, click + Configurations to configure the following sections:

​​

​

​The Text-to-Speech configuration page appears.

​

Basic Details

  • Name (required) - Enter a unique name for the STT configuration.
  • Description - Provide a brief description for internal reference.
  • Language (required) - Select the language for transcription.
    This determines the available providers and models.

AI Model Settings

This section defines the provider and model used to transcribe audio.

​

Provider (required) - Select a Speech-to-Text provider from the list. Depending on your account configuration, available providers may include:

  • Google Chirp

  • Sprinklr In-House

  • Eleven Labs

  • Azure Speech

  • Azure OpenAI

​

Model (provider-dependent)
The selected provider determines the underlying transcription model. Models may differ in:

  • Accuracy

  • Supported languages

  • Latency

  • Domain specialization

    ​

When Provider is set to Sprinklr In-House, you can select from the following available transcription models. The model list displays only supported transcription models registered for your environment.

  • Sprinklr AI Whisper

  • Wav2vec2

  • Whisper

​

Note: Contact Sprinklr Support to change or add any In-House model.

System Prompt (if supported by the provider)

Provide additional context to improve transcription quality.

For example:

  • Industry-specific terminology

  • Expected phrases or formats

  • Context about the conversation domain

​

Tip: Use system prompts when working with specialized domains such as banking, healthcare, or technical support.

​

Keyword Lists (Keyword Boosting)

Use Keyword Boosting to improve recognition of domain-specific terms.

  • Default Keyword List - Select a predefined keyword list.

Keyword lists help the system correctly recognize:

  • Brand names
  • Product names
  • Industry-specific vocabulary
  • Case linked variables
  • Profile linked variables
  • User level keywords (example, name of customer)

Example: Adding product names or company-specific terms improves transcription accuracy in customer conversations.

Multi Lingual Speech-To-Text

The Enable Multilingual Speech-to-Text Conversion option allows a single Speech-to-Text deployment to process conversations that may contain multiple languages.

When enabled, you can:

  1. Select a primary language for transcription.
  2. Add one or more additional languages supported by the selected provider and model.
  3. Associate a keyword boosting list with each language.
  4. Configure language-specific settings, such as accents, where supported by the selected provider.

The system dynamically displays available languages and supported accents based on the selected provider and model. Invalid language or accent combinations cannot be saved.

Example

A contact center serving customers in India can create a single STT deployment with:

  • Primary language: English
  • Additional language: Hindi
  • English keyword boosting list
  • Hindi keyword boosting list

During a conversation, based on the language passed, the configuration switches languages and applies the corresponding keyword boosting configuration to improve transcription quality.

​

Note: The number of supported languages depends on the capabilities and restrictions of the selected STT provider and model. Available languages, accents, and limits are populated dynamically during configuration.

Noise Suppression

Enable Noise Suppression to reduce background noise in incoming audio before transcription.

Noise Suppression can help improve speech recognition accuracy in environments where callers may experience:

  • Office background noise
  • Public environments
  • Call center ambient noise
  • Multiple speakers in close proximity
  • General environmental disturbances

Configure Noise Suppression

Field

Description

Model

Select the noise suppression model used to filter incoming audio. Currently, the In House Sprinklr model is supported.

Noise Suppression Threshold

Determines the level of noise reduction applied to incoming audio. Higher values provide stronger noise filtering, while lower values preserve more original audio.

Transcription Controls

This section defines how the final transcript is formatted before it is consumed by downstream systems.

Text Normalization

Enable Text Normalization to standardize transcription output.

This ensures consistent formatting of:

  • Dates
  • Numbers
  • Email addresses
  • Structured phrases

Example:

“twenty fifth May twenty twenty six” → “05/25/2026”

​

Word Replacements

Enable Word Replacements to substitute commonly mistranscribed words with preferred alternatives.

Use this to:

  • Correct recurring transcription errors
  • Standardize terminology

Example:
“sprinkler” → “Sprinklr”

​

Regex Replacement

Enable Regex Replacement to apply pattern-based transformations to structured text.

Use this for:

  • Ticket numbers
  • Account IDs
  • Reference codes

These rules are applied after transcription to improve data consistency.

Preview Configuration

Use the Preview Configuration panel (right side of the screen) to test the STT setup.

Audio Sample

  • Record or upload an audio sample
  • Optionally select Additional Hints for Testing to guide the system

Transcript Preview

  • View the generated transcript after processing
  • Validate accuracy and formatting

Note: If no sample is recorded, the transcript preview remains empty.

Where This Configuration Is Used

Once created, the STT configuration is available in Speech Profiles under:

Voice Care → Voice Settings → Speech Profiles → STT Source

Important Notes

  • STT configurations are reusable across deployments
  • Changes to a configuration affect all linked speech profiles
  • Available providers and features depend on language selection

​

Availability Note

STT Configuration is available only for new deployments created after this feature release. Existing deployments are not automatically migrated.

​

Related Articles

​