Kenpath Labs

Input streaming

Send text over a WebSocket as it is produced and receive audio continuously.

When to use it

Use this API when the text is still being produced — typically an LLM’s token stream. Rather than buffering output into sentences and synthesizing one at a time, the server consumes fragments as they arrive and holds back only a few words: the model generates each audio chunk from everything already spoken plus a short lookahead, so it never needs a complete sentence.

  • Speech starts 8 words into the stream with the default settings, mid-sentence, rather than waiting for the first sentence to end: ~80 ms to first audio
  • Prosody is continuous across chunk seams — each chunk is generated with the preceding audio as context, so there are no restarts or intonation resets at chunk boundaries
  • One connection carries one utterance, and it can be opened before the text exists — while the user is still speaking — so DNS, TLS and admission are off the critical path. Opening the socket ahead of time (the Python SDK’s speech.prepare()) is the fastest path to first audio

Two WebSocket endpoints expose this: a native one (recommended) and an ElevenLabs-compatible one for existing EL integrations. Both authenticate with your usual header on the upgrade request; see Authentication.

Native WebSocket

WS/v1/audio/speech/stream-input

Connect with query parameters, send text fragments as JSON messages the moment they’re produced, and read binary frames back: raw PCM, 16-bit LE mono at 24 kHz. The server chunks at safe boundaries and keeps rolling text-and-audio context, so chunk N+1 continues chunk N’s prosody. Don’t pre-chunk or sentence-split yourself.

ParameterTypeDefaultDescription
voicequery-Voice id from GET /v1/voices: an sv_-prefixed id, e.g. sv_enhdbrj5 (Aanya).
langquery-Language hint: any form works (hi, hin, hindi, hi-IN, ja, zh-CN, ko). Omit to auto-detect from script.
modequerysentenceChunking strategy: sentence (waits for sentence boundaries) or eager (word-level, lowest latency; see below).
speedquery1.0Speaking speed, 0.7–1.5, pitch-preserving. Applied continuously across the whole session, so chunk seams stay seamless.
sample_ratequery24000PCM rate of the frames you get back: 8000 16000 22050 24000 32000 44100 48000, the same set as HTTP. Resampled in-process in one stateful stream, so chunk boundaries carry no artifacts.
pronunciation_dictionary_idquery-Apply a pronunciation dictionary for the whole session.
temperature …query-Sampling knobs ride as query params too. Omit them; the server fills serving-tuned defaults.

Messages you send

ParameterTypeDefaultDescription
{"text": "fragment "}JSON-A text fragment, as small as a single token delta. Send them as fast as they arrive.
{"flush": true}JSON-Force everything buffered out as audio now (for example at the end of an LLM paragraph).
{"text": ""}JSON-End of stream. The server synthesizes the remainder, sends {"type": "done"}, and closes.

Messages you receive

ParameterTypeDefaultDescription
binary framebytes-PCM, 16-bit LE mono at sample_rate. Play it as it arrives.
{"type": "chunk", "text", "peek"}JSON-Sent before each chunk’s audio: the words about to be spoken and the lookahead they were generated with. Useful for captions and barge-in bookkeeping.
{"type": "flushed"}JSON-A flush you asked for has been spoken.
{"type": "done"}JSON-The utterance is complete; the server closes with code 1000. A socket that closes without it was cut short — treat the audio as truncated.
{"type": "error", "message"}JSON-The request was rejected (unknown voice, unsupported sample_rate, unparseable speed); the server then closes with 1008. Other close codes: 1013 the server is not ready, retry shortly; 1011 internal error.
pip install svara-voice
import asyncio
from openai import AsyncOpenAI
from svara import AsyncSvara
llm = AsyncOpenAI()
svara = AsyncSvara() # reads SVARA_API_KEY
async def speak(prompt: str):
async def deltas():
stream = await llm.chat.completions.create(
model="gpt-4o-mini", stream=True,
messages=[{"role": "user", "content": prompt}])
async for event in stream:
if event.choices[0].delta.content:
yield event.choices[0].delta.content
# One WebSocket; audio starts a few words into the LLM's output.
async for audio in svara.speech.stream_input(deltas(), voice="sv_enhdbrj5"):
player.feed(audio) # PCM16 @ 24 kHz mono
asyncio.run(speak("Explain how rainbows form, in three sentences."))

Eager mode

mode=eager is the lowest-latency configuration and maps directly onto how the model was trained. The server buffers incoming deltas at word granularity (mid-word token deltas are handled). Once enough complete words exist it synthesizes the first chunk_words with the next few words attached as the lookahead peek. Measured against production, first audio needs max(2 × chunk_words, chunk_words + peek_words) words:

Chinese and Japanese are written without spaces, so there the server counts one word per two characters and prefers to cut at 。!?,、.

ParameterTypeDefaultDescription
chunk_wordsquery4Words per synthesis chunk. Smaller = earlier first audio, more chunk seams (they’re seamless, but each chunk has scheduling overhead). 4 is the floor; lower values are raised to it.
peek_wordsquery2Lookahead words held back so each chunk knows what’s coming (clamped 1–5). The final chunk goes out with no peek, which is the model’s end-of-utterance signal.
max_chunk_wordsquery20Upper bound on a chunk once text has queued up — keeps a fast LLM from producing one enormous chunk.

With defaults, speech starts 8 words into the LLM’s output (9 with peek_words=5); typically well before the first sentence ends. An idle socket is not reaped on a short timer — one left open for five minutes still synthesized normally — so opening early is safe.

ElevenLabs-compatible WebSocket

WS/v1/text-to-speech/{voice_id}/stream-input

A drop-in implementation of the ElevenLabs realtime protocol, for existing EL integrations: the official EL SDKs’ WebSocket clients work unmodified. The full protocol is supported:

  • BOS: {"text": " ", "generation_config": {"chunk_length_schedule": [120,160,250,290]}} (schedule values clamped 50-500; auto_mode supported)
  • Text: {"text": "fragment ", "try_trigger_generation": true}, plus {"flush": true}; a bare space is a keepalive
  • EOS: {"text": ""}
  • Audio frames: {"audio": "<base64>", "alignment": …, "normalizedAlignment": …} in the negotiated output_format (query param, e.g. pcm_24000, mp3_44100_128), then {"isFinal": true} and a clean close (code 1000)
  • inactivity_timeoutquery param: 20 s default, 180 s max

Character timings in alignment are chunk-relative and approximate: chunk boundaries are sample-exact, characters uniform within a chunk.

Tips

  • Forward raw deltas. Don’t sentence-split, batch, or “clean up” the LLM stream client-side; the server’s chunker is the one that knows the model’s trained format.
  • Buffer ~150 ms on playback. A small jitter buffer on your audio sink absorbs network variance without hurting perceived latency.
  • Keep the socket warm for turns, not sessions. Open the connection when a reply starts, close on {"type": "done"}. Each connection is one utterance with one prosodic arc.
  • Concurrent conversations need concurrent streams: each open socket counts against your plan’s concurrency limit (see Rate limits).