Kenpath Labs

Text to speech

POST /v1/audio/speech: parameters, output formats, and language handling.

Endpoint

POST/v1/audio/speech

JSON in, audio bytes out. One request shape covers plain synthesis, streaming, and every output format. The response body is the raw audio in your requested response_format (or a chunked stream when stream: true).

Feeding text from an LLM, or holding a conversation? Use the WebSocket input-streaming API instead of per-request HTTP — one connection, speech that starts before the sentence ends.

Request parameters

input and voice are required. Everything else has a serving-tuned default: omit any parameter you don’t explicitly need rather than hardcoding a value, so your integration inherits improvements automatically.

ParameterTypeDefaultDescription
inputstringrequiredText to synthesize (max 5,000 chars). Raw graphemes in any supported language; code-switching within one string is fine. No SSML or phonemes.
voicestringrequiredLibrary voice id from GET /v1/voices: an sv_-prefixed id, e.g. sv_enhdbrj5 (Aanya).
modelstringsvara-tts-turboOptional. svara-tts-turbo is the one model the API serves (GET /v1/models). The field exists for OpenAI-SDK compatibility: any value is accepted and generation always uses the current model.
response_formatenumwavOne of mp3 opus aac flac wav pcm ulaw alaw. See Output formats below.
sample_rateint24000Output rate in Hz: 8000 16000 22050 24000 32000 44100 48000. Codec is 24 kHz native; others resampled in-process. opus is rendered at 8000, 16000, 24000 or 48000 only. The default applies to every format, ulaw and alaw included — ask for 8000 for telephony.
speedfloat1.0Speaking speed, 0.7–1.5. Pitch is preserved — only the pace changes. Works on streaming responses too.
streamboolfalseStream the audio as it’s generated. See Streaming.
langstringautoLanguage hint in any form (hi, hin, hindi, hi-IN, ja, zh-CN, ko). Enables number/unit normalization. Omit to auto-detect from script.
bitrate_kbpsint-For lossy formats (mp3/opus/aac), 8–320. Defaults: mp3 128, opus 64, aac 96.
pronunciation_dictionary_idstring-Apply a pronunciation dictionary (its UUID) to this request. An id the server cannot find is not an error: the workspace’s global rules apply and the response carries X-Svara-Dictionary: miss.
temperature, top_p, top_k, min_p, repetition_penalty, presence_penaltyfloat-Sampling knobs. Defaults are mode-aware and serving-tuned; leave them unset unless you have a measured reason.
Using an OpenAI SDK? The extended fields (stream, lang, sample_rate, sampling knobs) ride in the SDK’s extra_body parameter; see SDKs.

Output formats

ParameterTypeDefaultDescription
mp3lossy-Universal, small. Default 128 kbps. Best for web delivery and storage.
opuslossy-Efficient at low bitrates; great for real-time voice and WebRTC.
aaclossy-Apple-ecosystem friendly.
flaclossless-Archival / further processing without generation loss.
wavpcm container-Uncompressed, widely readable. The default.
pcmraw-Headerless 16-bit LE mono at your sample_rate. Lowest-latency streaming: feed straight to an audio sink.
ulaw / alawtelephony-8-bit companded for telephony; pair with sample_rate 8000.

Languages & normalization

The model covers 82 languages and switches between them mid-sentence. Send lang when you know it (any ISO form works) and the server normalizes numbers, dates, currency, and units into spoken form for that language. When you don’t know it, send nothing; the script is auto-detected. Don’t expose a “normalize” toggle in your UI; leave it on and let the server own it. Drive language pickers from GET /v1/languages.

Examples

curl -X POST https://api.kenpathlabs.com/v1/audio/speech \
-H "Authorization: Bearer $SVARA_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"voice": "sv_r22w7pwe",
"input": "Your appointment is confirmed for 3 PM tomorrow.",
"response_format": "ulaw",
"sample_rate": 8000
}' --output prompt.ulaw # voice sv_r22w7pwe = Aarav