Text to speech
POST /v1/audio/speech: parameters, output formats, and language handling.
Endpoint
/v1/audio/speechJSON in, audio bytes out. One request shape covers plain synthesis, streaming, and every output format. The response body is the raw audio in your requested response_format (or a chunked stream when stream: true).
Request parameters
input and voice are required. Everything else has a serving-tuned default: omit any parameter you don’t explicitly need rather than hardcoding a value, so your integration inherits improvements automatically.
| Parameter | Type | Default | Description |
|---|---|---|---|
| input | string | required | Text to synthesize (max 5,000 chars). Raw graphemes in any supported language; code-switching within one string is fine. No SSML or phonemes. |
| voice | string | required | Library voice id from GET /v1/voices: an sv_-prefixed id, e.g. sv_enhdbrj5 (Aanya). |
| model | string | svara-tts-turbo | Optional. svara-tts-turbo is the one model the API serves (GET /v1/models). The field exists for OpenAI-SDK compatibility: any value is accepted and generation always uses the current model. |
| response_format | enum | wav | One of mp3 opus aac flac wav pcm ulaw alaw. See Output formats below. |
| sample_rate | int | 24000 | Output rate in Hz: 8000 16000 22050 24000 32000 44100 48000. Codec is 24 kHz native; others resampled in-process. opus is rendered at 8000, 16000, 24000 or 48000 only. The default applies to every format, ulaw and alaw included — ask for 8000 for telephony. |
| speed | float | 1.0 | Speaking speed, 0.7–1.5. Pitch is preserved — only the pace changes. Works on streaming responses too. |
| stream | bool | false | Stream the audio as it’s generated. See Streaming. |
| lang | string | auto | Language hint in any form (hi, hin, hindi, hi-IN, ja, zh-CN, ko). Enables number/unit normalization. Omit to auto-detect from script. |
| bitrate_kbps | int | - | For lossy formats (mp3/opus/aac), 8–320. Defaults: mp3 128, opus 64, aac 96. |
| pronunciation_dictionary_id | string | - | Apply a pronunciation dictionary (its UUID) to this request. An id the server cannot find is not an error: the workspace’s global rules apply and the response carries X-Svara-Dictionary: miss. |
| temperature, top_p, top_k, min_p, repetition_penalty, presence_penalty | float | - | Sampling knobs. Defaults are mode-aware and serving-tuned; leave them unset unless you have a measured reason. |
stream, lang, sample_rate, sampling knobs) ride in the SDK’s extra_body parameter; see SDKs.Output formats
| Parameter | Type | Default | Description |
|---|---|---|---|
| mp3 | lossy | - | Universal, small. Default 128 kbps. Best for web delivery and storage. |
| opus | lossy | - | Efficient at low bitrates; great for real-time voice and WebRTC. |
| aac | lossy | - | Apple-ecosystem friendly. |
| flac | lossless | - | Archival / further processing without generation loss. |
| wav | pcm container | - | Uncompressed, widely readable. The default. |
| pcm | raw | - | Headerless 16-bit LE mono at your sample_rate. Lowest-latency streaming: feed straight to an audio sink. |
| ulaw / alaw | telephony | - | 8-bit companded for telephony; pair with sample_rate 8000. |
Languages & normalization
The model covers 82 languages and switches between them mid-sentence. Send lang when you know it (any ISO form works) and the server normalizes numbers, dates, currency, and units into spoken form for that language. When you don’t know it, send nothing; the script is auto-detected. Don’t expose a “normalize” toggle in your UI; leave it on and let the server own it. Drive language pickers from GET /v1/languages.