Overview
SarvamTTSService provides text-to-speech synthesis specialized for Indian languages and voices. The service uses the bulbul:v3 model by default, which offers temperature control for output randomness, pace adjustment, and 21 speaker voices. The previous bulbul:v2 model is deprecated as Sarvam’s API no longer serves it.
Sarvam TTS API Reference
Pipecat’s API methods for Sarvam AI TTS integration
Example Implementation
Complete example with Indian language support
Sarvam Documentation
Official Sarvam AI text-to-speech API documentation
Sarvam Console
Access Indian language voices and API keys
Installation
To use Sarvam AI services, no additional dependencies are required beyond the base installation:Prerequisites
Sarvam AI Account Setup
Before using Sarvam AI TTS services, you need:- Sarvam AI Account: Sign up at Sarvam AI Console
- API Key: Generate an API key from your account dashboard
- Language Selection: Choose from available Indian language voices
Required Environment Variables
SARVAM_API_KEY: Your Sarvam AI API key for authentication
Configuration
Sarvam offers two service implementations:SarvamTTSService (WebSocket) for real-time streaming and SarvamHttpTTSService (HTTP) for simpler batch synthesis.
SarvamTTSService
str
required
Sarvam AI API subscription key.
str
default:"bulbul:v3"
deprecated
TTS model to use. Options:
bulbul:v3, bulbul:v3-beta, bulbul:v2
(deprecated). Deprecated in v0.0.105. Use
settings=SarvamTTSService.Settings(model=...) instead.str
default:"None"
deprecated
Speaker voice ID. If
None, uses shubh (the default for v3). Deprecated in
v0.0.105. Use settings=SarvamTTSService.Settings(voice=...) instead.str
default:"wss://api.sarvam.ai/text-to-speech/ws"
WebSocket URL for the TTS backend.
TextAggregationMode
default:"TextAggregationMode.SENTENCE"
Controls how incoming text is aggregated before synthesis.
SENTENCE
(default) buffers text until sentence boundaries, producing more natural
speech. TOKEN streams tokens directly for lower latency. Import from
pipecat.services.tts_service.bool
default:"None"
deprecated
Deprecated in v0.0.104. Use
text_aggregation_mode instead.int
default:"None"
Audio sample rate in Hz (8000, 16000, 22050, 24000, and for v3: 32000, 44100,
48000). If
None, uses model-specific default (24000 for v3, 22050 for
deprecated v2).InputParams
default:"None"
deprecated
Deprecated in v0.0.105. Use
settings=SarvamTTSService.Settings(...)
instead.SarvamTTSService.Settings
default:"None"
Runtime-configurable settings. See SarvamTTSService
Settings below.
SarvamHttpTTSService
str
required
Sarvam AI API subscription key.
aiohttp.ClientSession
required
An aiohttp session for HTTP requests.
str
default:"bulbul:v3"
deprecated
TTS model to use. Options:
bulbul:v3, bulbul:v3-beta, bulbul:v2
(deprecated). Deprecated in v0.0.105. Use
settings=SarvamHttpTTSService.Settings(model=...) instead.str
default:"None"
deprecated
Speaker voice ID. If
None, uses shubh (the default for v3). Deprecated in
v0.0.105. Use settings=SarvamHttpTTSService.Settings(voice=...) instead.str
default:"https://api.sarvam.ai"
Sarvam AI API base URL.
int
default:"None"
Audio sample rate in Hz (8000, 16000, 22050, 24000, and for v3: 32000, 44100,
48000). If
None, uses model-specific default (24000 for v3, 22050 for
deprecated v2).InputParams
default:"None"
deprecated
Deprecated in v0.0.105. Use
settings=SarvamHttpTTSService.Settings(...)
instead.SarvamHttpTTSService.Settings
default:"None"
Runtime-configurable settings. See SarvamHttpTTSService
Settings below.
SarvamTTSService Settings
Runtime-configurable settings passed via thesettings constructor argument using SarvamTTSService.Settings(...). These can be updated mid-conversation with TTSUpdateSettingsFrame. See Service Settings for details.
SarvamHttpTTSService Settings
Runtime-configurable settings passed via thesettings constructor argument using SarvamHttpTTSService.Settings(...). These can be updated mid-conversation with TTSUpdateSettingsFrame. See Service Settings for details.
Usage
Basic Setup (WebSocket)
With Temperature Control
HTTP Service
Notes
- Model default:
bulbul:v3is the default model. The previousbulbul:v2is deprecated since 1.9.0 — Sarvam’s API rejects it. - Model differences:
bulbul:v2supported pitch and loudness control;bulbul:v3supports temperature control but does not support pitch or loudness. Setting unsupported parameters for a model will log a warning. - Default voice: v3 defaults to
shubh; deprecated v2 defaulted toanushka. - Default sample rate: v3 defaults to 24000 Hz; deprecated v2 defaulted to 22050 Hz.
- Indian language focus: Sarvam AI specializes in Indian languages, supporting Bengali, English (India), Gujarati, Hindi, Kannada, Malayalam, Marathi, Odia, Punjabi, Tamil, and Telugu.
- Pace ranges differ:
bulbul:v3supports pace from 0.5 to 2.0; deprecated v2 supported 0.3 to 3.0. Values outside the range are clamped automatically.