Enhance Sprinklr Voice Bots with Speech-to-Speech (S2S) Models
Updated
Sprinklr supports real-time Speech-to-Speech (S2S) models through the OpenAI Realtime API. With S2S, Voice Bots can process spoken input and generate spoken responses directly, reducing latency and creating more natural conversations.
Currently, S2S supports a limited set of Voice Bot capabilities. Additional features will be introduced in future releases. This article explains how S2S works in Sprinklr and describes how to enable it for a customer environment.
How Speech-to-Speech Works in Sprinklr
Sprinklr integrates with the OpenAI Realtime API to support Speech-to-Speech interactions. When S2S is enabled, the Voice Bot processes audio input and output through the S2S model, eliminating the need for separate speech recognition and text-to-speech processing. This approach helps deliver faster and more natural conversations.
Before you can use S2S, the model must be configured for the customer environment. Contact the Conversational AI team for assistance with onboarding and configuration.
Enable Speech-to-Speech Models
After the model has been onboarded, complete the following steps.
Request Feature Enablement
Submit a support request to enable Speech-to-Speech for the customer environment.
Enable Speech-to-Speech in the IVR Flow
- Open the Redirect to Voice Bot node in the IVR flow.
- Turn on the Enable Speech-to-Speech Model toggle.
This setting activates the real-time S2S processing pipeline for the selected Voice Bot.

Select the Voice-Enabled AI Agent
- In the Redirect to Voice Bot node, open the Voice Bot drop-down list.
- Select the appropriate Voice-Enabled AI Agent application.
This ensures that the Voice Bot uses the configured Speech-to-Speech model for real-time voice interactions.