# SDKs & compatibility

The Svara Python SDK, agent skills for coding agents, and using the OpenAI or ElevenLabs SDKs against Svara.

## Svara Python SDK

The official SDK. One dependency-light package (`httpx` + `websockets`) that covers the whole surface: sync and async clients, HTTP streaming, the [input-streaming WebSocket](https://docs.kenpathlabs.com/input-streaming.md) behind one method call (with a blocking twin for code without an event loop), character timestamps, pronunciation dictionaries, usage, a CLI with a connectivity check (`svara doctor`), and drop-in [LiveKit](https://docs.kenpathlabs.com/livekit.md) / [Pipecat](https://docs.kenpathlabs.com/pipecat.md) plugins. Its transport defaults are measured against production for voice agents: connections are kept between turns, and the streaming socket can be opened before the text exists. The distribution is `svara-voice`; the import is `import svara`.

```bash
pip install svara-voice
```

```python
from svara import Svara

client = Svara()                          # reads SVARA_API_KEY
audio = client.speech.create(
    input="नमस्ते! Welcome to Svara.",
    voice="sv_enhdbrj5",                  # Aanya
    response_format="mp3",
)
audio.save("hello.mp3")                   # bytes, plus .sample_rate, .rate_limit, .request_id

# low latency: stream chunks as they generate (pcm by default)
stream = client.speech.stream(input="...", voice="sv_enhdbrj5")
for chunk in stream:
    player.feed(chunk)
print(stream.time_to_first_audio)         # seconds, measured at your client
```

```python
from svara import AsyncSvara

client = AsyncSvara()

# the endpoint voice agents should be on: feed an LLM token stream,
# audio starts before the sentence ends (one WebSocket, not N requests)
async for audio in client.speech.stream_input(llm_deltas(), voice="sv_enhdbrj5"):
    player.feed(audio)

# fastest: open the socket while the user is still talking
prepared = await client.speech.prepare(voice="sv_enhdbrj5")
async for audio in prepared.stream(llm_deltas()):
    player.feed(audio)
```

```
svara say "नमस्ते दुनिया" --voice sv_enhdbrj5 --out hello.mp3
svara voices --language hi
svara usage
svara doctor          # DNS, TLS, key, HTTP and WebSocket synthesis, with timings
```

> Extras: `pip install "svara-voice[livekit]"` and `pip install "svara-voice[pipecat]"`. Moving from the OpenAI SDK? `client.audio.speech.create(…)`, `.write_to_file()` and `with_streaming_response` run on a `Svara` client as written — only the constructor changes. Source, changelog and the full reference are on [GitHub](https://github.com/kenpath-labs/svara-python).

## Agent skills

Svara publishes skills for coding agents. The list of skills and what each one covers are in the [skills folder on GitHub](https://github.com/kenpath-labs/svara-python/tree/main/skills). Install them with:

- Any coding agent (Claude Code, Codex, Cursor, Gemini CLI, Copilot): `npx skills add kenpath-labs/svara-python`
- Claude Code: `/plugin marketplace add kenpath-labs/svara-python`, then `/plugin install svara@kenpath-labs`

> The Python SDK and agent skills are open source (Apache-2.0). The Svara TTS Turbo model and API are proprietary to Kenpath Labs.

## OpenAI SDKs

The official OpenAI Python and JavaScript SDKs work against Svara unmodified: point `base_url` at Svara and pass your key. Svara-specific fields ride in `extra_body` (Python) or as extra properties (JS).

```bash
pip install openai
```

```python
from openai import OpenAI

client = OpenAI(base_url="https://api.kenpathlabs.com/v1", api_key=SVARA_API_KEY)
audio = client.audio.speech.create(
    model="svara-tts-turbo",
    voice="sv_enhdbrj5",  # Aanya
    input="Namaste!",
    response_format="mp3",
    extra_body={"lang": "hindi", "stream": False},   # svara extensions
)
audio.write_to_file("out.mp3")
```

```javascript
import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://api.kenpathlabs.com/v1", apiKey: SVARA_API_KEY });
const res = await client.audio.speech.create({
  model: "svara-tts-turbo",
  voice: "sv_enhdbrj5", // Aanya
  input: "Namaste!",
  response_format: "mp3",
  // @ts-expect-error - svara extension
  lang: "hindi",
});
```

## ElevenLabs SDKs

The official ElevenLabs SDKs also work unmodified, including their realtime WebSocket client. Pass any non-empty `api_key`; its presence is also how `GET /v1/models` decides to answer in the ElevenLabs array shape.

```bash
pip install elevenlabs
```

```python
from elevenlabs.client import ElevenLabs

client = ElevenLabs(base_url="https://api.kenpathlabs.com", api_key=SVARA_API_KEY)
audio = client.text_to_speech.convert(
    voice_id="sv_enhdbrj5",  # Aanya
    text="Namaste!",
    model_id="eleven_multilingual_v2",   # accepted, ignored
    output_format="mp3_44100_128",
)
```

```javascript
import { ElevenLabsClient } from "@elevenlabs/elevenlabs-js";

const client = new ElevenLabsClient({ baseUrl: "https://api.kenpathlabs.com", apiKey: SVARA_API_KEY });
const stream = await client.textToSpeech.stream("sv_enhdbrj5", { text: "Namaste!" }); // voice: Aanya
```

## Base URL conventions

> Mind the `/v1`: **OpenAI SDKs include it** in the base URL (`https://api.kenpathlabs.com/v1`), while **ElevenLabs SDKs do not** (`https://api.kenpathlabs.com`); they append `v1/…` themselves.

## Compatibility surface

What maps, and what to expect:

- **Full support**: TTS + streaming, all output formats and rates, timestamps, the realtime WebSocket, the voice list and search (`voices.get_all()`, `voices.search()`), single voice, models, languages. For LLM-driven speech prefer the native WebSocket: the ElevenLabs realtime protocol buffers ~120 characters before its first generation, so first audio comes noticeably later than the ~80 ms on the native socket.
- **Honored**: `voice_settings.speed` maps onto the native speed parameter (their 0.7–1.2 range sits inside our 0.7–1.5) — on HTTP, with-timestamps and the realtime WebSocket alike.
- **Accepted and ignored**: `voice_settings` (stability/similarity/style), `seed`, `model_id`, `previous_text`/`next_text`, request stitching. These don’t map onto this model; requests including them still succeed. (Pronunciation dictionaries _are_ supported — see [the guide](https://docs.kenpathlabs.com/pronunciation.md).)
- **Stubbed**: user/subscription and voice-settings endpoints return valid, permissive shapes so SDK flows don’t break.

A machine-readable OpenAPI document is served live at `https://api.kenpathlabs.com/openapi.json`: generate typed clients from it (Fern, Stainless, Scalar) for languages the official SDK doesn’t cover yet.
