zoxaAI
Homepage

TTS Providers

Every text-to-speech provider supported by zoxaAI -- voices, speed ranges, volume controls, and language support per provider.

Overview

zoxaAI supports 6 TTS providers (8 models). Each agent is configured with exactly one TTS provider. Voice selection, speed, and other synthesis parameters are provider-specific. Models are a curated catalog — you pick from the models listed per provider below.


Provider Comparison

ProviderDefault ModelSpeed RangeVolumeLanguagesVoice Selection
ElevenLabseleven_flash_v2_50.7 - 1.2No12 (flash) / 21 (v3) / 22 (v4 Turbo)Voice ID from ElevenLabs library
Cartesiasonic-3.60.6 - 1.50.5 - 2.0Via voice/modelVoice UUID from Cartesia library
Smallestlightning_v3.10.5 - 2.0No16 languagesNamed voices + custom
Sarvambulbul:v30.5 - 2.0No11 Indian locales38 voices, temperature knob (v3)
xAIxai-tts0.7 - 1.5No12 languages5 built-in voices
Inworldinworld-tts-20.5 - 1.5No22 languages269-voice catalog + custom

ElevenLabs

Default voice: 7qBNUtXRGP0jPi0H4r8k

Models

ModelNotes
eleven_v4_turboDefault -- most expressive ElevenLabs model (Eleven v4 Turbo) — performs inline audio tags like [laughs] and [excited], covers every platform language. Reads only the stability voice setting; the other voice knobs don't apply.
eleven_v3_conversationalEleven v3 — expressive, performs inline audio tags, covers every platform language except Odia. Higher latency (~334 ms) than v4 Turbo. Same Text-to-Dialogue connection as v4 Turbo: reads only stability.
eleven_flash_v2_5Best balance of quality and speed, lowest latency (~75 ms). The only ElevenLabs model that reads every voice setting below, including speed.

Configuration Fields

FieldTypeDefaultRangeDescription
voicestring7qBNUtXRGP0jPi0H4r8k--ElevenLabs voice ID. Paste any voice ID from the ElevenLabs voice library.
modelstringeleven_v4_turbo--Synthesis model. Flash models have lower latency.
speedfloat1.00.7 - 1.2Speech speed multiplier.
stabilityfloat0.50.0 - 1.0Voice stability. Lower values add more expressiveness and variation.
similarity_boostfloat0.750.0 - 1.0Voice similarity. Higher values make output closer to the original voice sample.
stylefloat0.00.0 - 1.0Style exaggeration. Higher values amplify the voice's style. Increases latency.
use_speaker_boostbooltrue--Boost similarity to the original speaker.
apply_text_normalizationenumautoauto | on | offSpeak numbers, dates, and abbreviations naturally.

Any voice ID from the ElevenLabs voice library is accepted — paste it directly into the voice field. You can verify custom voice IDs through the dashboard voice selector.

Language support: Comes from the agent's ordered languages list (languages[0] = primary) — there is no per-provider language field. Coverage is per model: Flash v2.5 covers English, Chinese, Spanish, Arabic, French, Portuguese, Russian, Indonesian, German, Japanese, Hindi, and Tamil; Eleven v3 covers every platform language except Odia; Eleven v4 Turbo covers every platform language. The agent's language picker greys a model out for anything outside its set.

Increasing the style parameter above 0.0 adds latency to every TTS request. For voice pipelines where speed is critical, keep style at 0.0 unless expressiveness is a priority.


Cartesia

Default model: sonic-3.6

Models

ModelNotes
sonic-3.6Latest, highest quality, 21 of the platform languages including Odia and Urdu. The only Cartesia model exposed by zoxaAI.

Configuration Fields

FieldTypeDefaultRangeDescription
voicestringc63361f8-d142-4c62-8da7-8f8149d973d6--Cartesia voice UUID (default is the platform's Krishna voice). Browse voices in the Cartesia voice library.
modelstringsonic-3.6--Synthesis model.
speedfloat1.00.6 - 1.5Speech speed multiplier.
volumefloat1.50.5 - 2.0Volume multiplier for generated speech — 1.0 leaves the voice unchanged; the platform default is 1.5.

Cartesia is the only TTS provider that supports volume control, and it pairs that with speed control. This makes it a good choice when you need fine-grained control over speech output characteristics.

Language support: Multilingual -- language is controlled by the voice selected and the model. Cartesia does not expose a separate language configuration field.


Smallest

Default model: lightning_v3.1

Models

ModelNotes
lightning_v3.1Realtime model with the full lightning-v3.1 voice catalog. The Smallest model exposed by zoxaAI.

Configuration Fields

FieldTypeDefaultRangeDescription
voicestringsophia--Voice name. Custom input accepted -- paste any voice_id from Smallest's full catalog.
modelstringlightning_v3.1--Synthesis model.
speedfloat1.00.5 - 2.0Speech speed multiplier.

Language selection comes from the agent's ordered languages list (languages[0] = primary) — there is no per-provider language field.

Curated Voices

The dropdown surfaces a focused subset of lightning_v3.1's ~240-voice catalog (American English, British English, and Hinglish). For any other voice, paste its voice_id directly.

VoiceAccentGender
sophia (default), avery, mia, christineAmerican EnglishFemale
alex, ethan, robertAmerican EnglishMale
poppyBritish EnglishFemale
liam, noahBritish EnglishMale
maya, aisha, advikaHinglish (Indian)Female
arjun, vivaan, kaustubhHinglish (Indian)Male

Supported Languages

CodeLanguageCodeLanguage
enEnglishesSpanish
hiHindifrFrench
bnBengalideGerman
guGujaratiptPortuguese
knKannadaruRussian
mlMalayalam
mrMarathi
orOdia
paPunjabi
taTamil
teTelugu

Smallest has strong support for Indian languages (Hindi, Bengali, Gujarati, Kannada, Malayalam, Marathi, Odia, Punjabi, Tamil, Telugu), making it a good complement to Sarvam for Indian language deployments.


xAI

Default model: xai-tts (fixed)

Configuration Fields

FieldTypeDefaultRangeDescription
voicestringeve--xAI built-in voice name (one of the 5 below).
modelstringxai-tts--Fixed model identifier.
speedfloat1.00.7 - 1.5Speaking rate.
optimizeStreamingLatencyint00 - 2Latency mode: 0 = best quality, 2 = lowest time-to-first-audio.
textNormalizationboolfalse--Speak numbers and dates naturally — adds a little latency.

Available Voices

VoiceNotes
eveDefault
ara
leo
rex
sal

Supported Languages

en, zh, es, ar, fr, pt, ru, id, de, ja, hi, bn (12 languages).

xAI TTS streams over a WebSocket at the pipeline's 24 kHz. You can try inline markup like [laugh], [pause], and <whisper> in the text. Timestamps are char-level.


Sarvam

Default model: bulbul:v3

Models

ModelNotes
bulbul:v3Default -- 38 voices

Configuration Fields

FieldTypeDefaultRangeDescription
voicestringishita--Voice name from the bulbul:v3 catalog.
modelstringbulbul:v3--Synthesis model.
pacefloat1.00.5 - 2.0Speech speed multiplier. Sent to Sarvam as the native pace parameter.
temperaturefloat0.60.01 - 1.0Bulbul v3 only. Output randomness / expressiveness. Lower values produce a steadier, more predictable read; higher values add variation between takes.

Language selection comes from the agent's ordered languages list (languages[0] = primary) — there is no per-provider language field. The factory maps each selected language to Sarvam's xx-IN code.

Voices

38 voices: shubh, aditya, ritu, priya, neha, rahul, pooja, rohan, simran, kavya, amit, dev, ishita, shreya, ratan, varun, manan, sumit, roopa, kabir, aayan, ashutosh, advait, amelia, sophia, anand, tanya, tarun, sunny, mani, gokul, vijay, shruti, suhani, mohit, kavitha, rehan, soham, rupali

Supported Languages

CodeLanguageCodeLanguage
en-INEnglish (India)mr-INMarathi
hi-INHindior-INOdia
bn-INBengalipa-INPunjabi
gu-INGujaratita-INTamil
kn-INKannadate-INTelugu
ml-INMalayalam

Sarvam specializes in Indian languages (11 Indian locales). Language codes use the xx-IN format.


Inworld

Default model: inworld-tts-2

Models

ModelNotes
inworld-tts-2Default — supports all 22 languages.
inworld-tts-2-flashFastest, lowest-cost Inworld model ($15 per 1M characters vs $25 for tts-2) — same 22 languages and voice catalog.

Configuration Fields

FieldTypeDefaultRangeDescription
voicestringAshley--Inworld voice name. Pick from the full 269-voice catalog in the voice picker, or paste any custom voice_id from the Inworld studio.
modelstringinworld-tts-2--Synthesis model.
speakingRatefloat1.00.5 - 1.5Speaking rate.
temperaturefloat1.00.01 - 2.0Delivery variability. inworld-tts-2-flash only — Inworld ignores temperature on inworld-tts-2, so the editor hides it for that model.

Voices

The full official Inworld catalog — 269 preset voices across 15 languages — is available in the voice picker, with search and gender filters. Ashley is the default. Any custom voice id cloned in the Inworld studio is also accepted via the picker's manual-entry toggle.

Supported Languages

  • inworld-tts-2 and inworld-tts-2-flash — all 22 platform languages.

Word-level timestamps are available in every supported language, on both models.


Choosing a Provider

On this page