TTS Providers
Every text-to-speech provider supported by zoxaAI -- voices, speed ranges, volume controls, and language support per provider.
Overview
zoxaAI supports 6 TTS providers (8 models). Each agent is configured with exactly one TTS provider. Voice selection, speed, and other synthesis parameters are provider-specific. Models are a curated catalog — you pick from the models listed per provider below.
Provider Comparison
| Provider | Default Model | Speed Range | Volume | Languages | Voice Selection |
|---|---|---|---|---|---|
| ElevenLabs | eleven_flash_v2_5 | 0.7 - 1.2 | No | 12 (flash) / 21 (v3) / 22 (v4 Turbo) | Voice ID from ElevenLabs library |
| Cartesia | sonic-3.6 | 0.6 - 1.5 | 0.5 - 2.0 | Via voice/model | Voice UUID from Cartesia library |
| Smallest | lightning_v3.1 | 0.5 - 2.0 | No | 16 languages | Named voices + custom |
| Sarvam | bulbul:v3 | 0.5 - 2.0 | No | 11 Indian locales | 38 voices, temperature knob (v3) |
| xAI | xai-tts | 0.7 - 1.5 | No | 12 languages | 5 built-in voices |
| Inworld | inworld-tts-2 | 0.5 - 1.5 | No | 22 languages | 269-voice catalog + custom |
ElevenLabs
Default voice: 7qBNUtXRGP0jPi0H4r8k
Models
| Model | Notes |
|---|---|
eleven_v4_turbo | Default -- most expressive ElevenLabs model (Eleven v4 Turbo) — performs inline audio tags like [laughs] and [excited], covers every platform language. Reads only the stability voice setting; the other voice knobs don't apply. |
eleven_v3_conversational | Eleven v3 — expressive, performs inline audio tags, covers every platform language except Odia. Higher latency (~334 ms) than v4 Turbo. Same Text-to-Dialogue connection as v4 Turbo: reads only stability. |
eleven_flash_v2_5 | Best balance of quality and speed, lowest latency (~75 ms). The only ElevenLabs model that reads every voice setting below, including speed. |
Configuration Fields
| Field | Type | Default | Range | Description |
|---|---|---|---|---|
voice | string | 7qBNUtXRGP0jPi0H4r8k | -- | ElevenLabs voice ID. Paste any voice ID from the ElevenLabs voice library. |
model | string | eleven_v4_turbo | -- | Synthesis model. Flash models have lower latency. |
speed | float | 1.0 | 0.7 - 1.2 | Speech speed multiplier. |
stability | float | 0.5 | 0.0 - 1.0 | Voice stability. Lower values add more expressiveness and variation. |
similarity_boost | float | 0.75 | 0.0 - 1.0 | Voice similarity. Higher values make output closer to the original voice sample. |
style | float | 0.0 | 0.0 - 1.0 | Style exaggeration. Higher values amplify the voice's style. Increases latency. |
use_speaker_boost | bool | true | -- | Boost similarity to the original speaker. |
apply_text_normalization | enum | auto | auto | on | off | Speak numbers, dates, and abbreviations naturally. |
Any voice ID from the ElevenLabs voice library is accepted — paste it directly into the voice field. You can verify custom voice IDs through the dashboard voice selector.
Language support: Comes from the agent's ordered languages list (languages[0] = primary) — there is no per-provider language field. Coverage is per model: Flash v2.5 covers English, Chinese, Spanish, Arabic, French, Portuguese, Russian, Indonesian, German, Japanese, Hindi, and Tamil; Eleven v3 covers every platform language except Odia; Eleven v4 Turbo covers every platform language. The agent's language picker greys a model out for anything outside its set.
Increasing the style parameter above 0.0 adds latency to every TTS request. For voice pipelines where speed is critical, keep style at 0.0 unless expressiveness is a priority.
Cartesia
Default model: sonic-3.6
Models
| Model | Notes |
|---|---|
sonic-3.6 | Latest, highest quality, 21 of the platform languages including Odia and Urdu. The only Cartesia model exposed by zoxaAI. |
Configuration Fields
| Field | Type | Default | Range | Description |
|---|---|---|---|---|
voice | string | c63361f8-d142-4c62-8da7-8f8149d973d6 | -- | Cartesia voice UUID (default is the platform's Krishna voice). Browse voices in the Cartesia voice library. |
model | string | sonic-3.6 | -- | Synthesis model. |
speed | float | 1.0 | 0.6 - 1.5 | Speech speed multiplier. |
volume | float | 1.5 | 0.5 - 2.0 | Volume multiplier for generated speech — 1.0 leaves the voice unchanged; the platform default is 1.5. |
Cartesia is the only TTS provider that supports volume control, and it pairs that with speed control. This makes it a good choice when you need fine-grained control over speech output characteristics.
Language support: Multilingual -- language is controlled by the voice selected and the model. Cartesia does not expose a separate language configuration field.
Smallest
Default model: lightning_v3.1
Models
| Model | Notes |
|---|---|
lightning_v3.1 | Realtime model with the full lightning-v3.1 voice catalog. The Smallest model exposed by zoxaAI. |
Configuration Fields
| Field | Type | Default | Range | Description |
|---|---|---|---|---|
voice | string | sophia | -- | Voice name. Custom input accepted -- paste any voice_id from Smallest's full catalog. |
model | string | lightning_v3.1 | -- | Synthesis model. |
speed | float | 1.0 | 0.5 - 2.0 | Speech speed multiplier. |
Language selection comes from the agent's ordered languages list (languages[0] = primary) — there is no per-provider language field.
Curated Voices
The dropdown surfaces a focused subset of lightning_v3.1's ~240-voice catalog (American English, British English, and Hinglish). For any other voice, paste its voice_id directly.
| Voice | Accent | Gender |
|---|---|---|
sophia (default), avery, mia, christine | American English | Female |
alex, ethan, robert | American English | Male |
poppy | British English | Female |
liam, noah | British English | Male |
maya, aisha, advika | Hinglish (Indian) | Female |
arjun, vivaan, kaustubh | Hinglish (Indian) | Male |
Supported Languages
| Code | Language | Code | Language | |
|---|---|---|---|---|
en | English | es | Spanish | |
hi | Hindi | fr | French | |
bn | Bengali | de | German | |
gu | Gujarati | pt | Portuguese | |
kn | Kannada | ru | Russian | |
ml | Malayalam | |||
mr | Marathi | |||
or | Odia | |||
pa | Punjabi | |||
ta | Tamil | |||
te | Telugu |
Smallest has strong support for Indian languages (Hindi, Bengali, Gujarati, Kannada, Malayalam, Marathi, Odia, Punjabi, Tamil, Telugu), making it a good complement to Sarvam for Indian language deployments.
xAI
Default model: xai-tts (fixed)
Configuration Fields
| Field | Type | Default | Range | Description |
|---|---|---|---|---|
voice | string | eve | -- | xAI built-in voice name (one of the 5 below). |
model | string | xai-tts | -- | Fixed model identifier. |
speed | float | 1.0 | 0.7 - 1.5 | Speaking rate. |
optimizeStreamingLatency | int | 0 | 0 - 2 | Latency mode: 0 = best quality, 2 = lowest time-to-first-audio. |
textNormalization | bool | false | -- | Speak numbers and dates naturally — adds a little latency. |
Available Voices
| Voice | Notes |
|---|---|
eve | Default |
ara | |
leo | |
rex | |
sal |
Supported Languages
en, zh, es, ar, fr, pt, ru, id, de, ja, hi, bn (12 languages).
xAI TTS streams over a WebSocket at the pipeline's 24 kHz. You can try inline markup like [laugh], [pause], and <whisper> in the text. Timestamps are char-level.
Sarvam
Default model: bulbul:v3
Models
| Model | Notes |
|---|---|
bulbul:v3 | Default -- 38 voices |
Configuration Fields
| Field | Type | Default | Range | Description |
|---|---|---|---|---|
voice | string | ishita | -- | Voice name from the bulbul:v3 catalog. |
model | string | bulbul:v3 | -- | Synthesis model. |
pace | float | 1.0 | 0.5 - 2.0 | Speech speed multiplier. Sent to Sarvam as the native pace parameter. |
temperature | float | 0.6 | 0.01 - 1.0 | Bulbul v3 only. Output randomness / expressiveness. Lower values produce a steadier, more predictable read; higher values add variation between takes. |
Language selection comes from the agent's ordered languages list (languages[0] = primary) — there is no per-provider language field. The factory maps each selected language to Sarvam's xx-IN code.
Voices
38 voices: shubh, aditya, ritu, priya, neha, rahul, pooja, rohan, simran, kavya, amit, dev, ishita, shreya, ratan, varun, manan, sumit, roopa, kabir, aayan, ashutosh, advait, amelia, sophia, anand, tanya, tarun, sunny, mani, gokul, vijay, shruti, suhani, mohit, kavitha, rehan, soham, rupali
Supported Languages
| Code | Language | Code | Language | |
|---|---|---|---|---|
en-IN | English (India) | mr-IN | Marathi | |
hi-IN | Hindi | or-IN | Odia | |
bn-IN | Bengali | pa-IN | Punjabi | |
gu-IN | Gujarati | ta-IN | Tamil | |
kn-IN | Kannada | te-IN | Telugu | |
ml-IN | Malayalam |
Sarvam specializes in Indian languages (11 Indian locales). Language codes use the xx-IN format.
Inworld
Default model: inworld-tts-2
Models
| Model | Notes |
|---|---|
inworld-tts-2 | Default — supports all 22 languages. |
inworld-tts-2-flash | Fastest, lowest-cost Inworld model ($15 per 1M characters vs $25 for tts-2) — same 22 languages and voice catalog. |
Configuration Fields
| Field | Type | Default | Range | Description |
|---|---|---|---|---|
voice | string | Ashley | -- | Inworld voice name. Pick from the full 269-voice catalog in the voice picker, or paste any custom voice_id from the Inworld studio. |
model | string | inworld-tts-2 | -- | Synthesis model. |
speakingRate | float | 1.0 | 0.5 - 1.5 | Speaking rate. |
temperature | float | 1.0 | 0.01 - 2.0 | Delivery variability. inworld-tts-2-flash only — Inworld ignores temperature on inworld-tts-2, so the editor hides it for that model. |
Voices
The full official Inworld catalog — 269 preset voices across 15 languages — is available in the voice picker, with search and gender filters. Ashley is the default. Any custom voice id cloned in the Inworld studio is also accepted via the picker's manual-entry toggle.
Supported Languages
inworld-tts-2andinworld-tts-2-flash— all 22 platform languages.
Word-level timestamps are available in every supported language, on both models.