STT Providers
Every speech-to-text provider supported by zoxaAI -- models, language support, and configuration options per provider.
Overview
zoxaAI supports 5 STT providers (6 models). Each agent is configured with exactly one STT provider. Language support, model selection, and recognition features vary by provider. Models are a curated catalog — you pick from the models listed per provider below.
Provider Comparison
| Provider | Default Model | Languages | Keyword Boosting | Notes |
|---|---|---|---|---|
| Soniox | stt-rt-v5 | 56+ | Yes (contextTerms) | Platform default. Native code-switching — the agent's whole languages list stays live in one stream, no model swap |
| Deepgram Flux | flux-general-en | English (flux-general-en) / 8 (flux-general-multi) | Yes (keyterms) | Server-side end-of-turn detection |
| AssemblyAI | universal-3-6-pro | 12 (of the model's 32) | Yes (keyterms) | High accuracy; latency/accuracy mode presets |
| Cartesia | ink-2 | 5 (auto-detected) | Yes (keyterms) | Server-side turn detection with tunable thresholds |
| Sarvam | saaras:v3 | 13 (Indian locales) + code-mix | No | Indian language specialist; server-side VAD |
Soniox
Default model: stt-rt-v5 — the platform default STT provider.
Models
| Model | Notes |
|---|---|
stt-rt-v5 | Default — single realtime streaming model (Soniox's stt-rt-v5). |
Configuration Fields
| Field | Type | Default | Range | Description |
|---|---|---|---|---|
model | string | stt-rt-v5 | -- | Recognition model. |
contextTerms | string | "" | -- | Comma-separated bias terms (names, jargon) to improve recognition. Empty = off. |
endpointSensitivity | float | 0.3 | -1.0 - 1.0 | How eagerly Soniox ends a turn. Higher = faster endpointing. |
endpointLatencyAdjustmentLevel | int | 2 | 0 - 3 | Latency-vs-accuracy tradeoff for endpoint detection. |
maxEndpointDelayMs | int | 2000 | 500 - 3000 | Upper bound on how long Soniox waits before forcing a turn end (platform-uniform 2s silence cap). |
Supported Languages
Soniox recognizes 56+ languages in a single realtime model. Language selection comes from the agent's ordered languages list — Soniox receives the whole list as live hints, including en, zh, es, ar, fr, pt, ru, id, de, ja, hi, bn, gu, kn, ml, mr, pa, ta, te, ur.
Why Soniox for multilingual phone calls
Most STTs that claim "multilingual" pick one language up front and recognize that one. Soniox keeps every language in the agent's languages list live in the same stream — the model expects code-switching mid-utterance and handles it natively. That makes it the best fit for Indian voice agents (English ↔ Hindi mid-sentence) and any market where speakers routinely mix two languages.
Example
{
"languages": ["en", "hi"],
"stt": {
"provider": "soniox",
"soniox": { "contextTerms": "Acme, zoxa" }
}
}Deepgram Flux
Default model: flux-general-en
Models
| Model | Description |
|---|---|
flux-general-en | Default -- English only. Hides the language picker. |
flux-general-multi | Multilingual (en, es, fr, pt, ru, de, ja, hi). |
Configuration Fields
| Field | Type | Default | Range | Description |
|---|---|---|---|---|
model | string | flux-general-en | -- | Recognition model. |
eotThreshold | float | 0.7 | 0.5 - 0.9 | End-of-turn confidence needed before the turn is finalized. |
eotTimeoutMs | int | 2000 | 500 - 60000 | Max wait before the turn is force-ended (platform-uniform 2s silence cap). |
keyterms | string | "" | -- | Comma-separated bias terms, hot-applied mid-stream by Flux. |
Flux runs server-side end-of-turn detection instead of a local silence timer. flux-general-en only supports English and hides the language picker; switch to flux-general-multi for the multilingual set. Both models support keyword boosting via keyterms.
AssemblyAI
Default model: universal-3-6-pro
Models
| Model | Description |
|---|---|
universal-3-6-pro | Default -- Universal-3.6 Pro Realtime: streams 32 languages with native code-switching; 12 of them are platform languages (below). |
Configuration Fields
| Field | Type | Default | Range | Description |
|---|---|---|---|---|
model | string | universal-3-6-pro | -- | Recognition model. |
mode | enum | balanced | balanced | min_latency | max_accuracy | Preset that shifts several server defaults at once — min_latency responds fastest, max_accuracy waits longest before ending a turn. |
keyterms | string | "" | -- | Comma-separated bias terms (up to 100 terms / 50 chars each). |
voiceFocus | enum | near-field | off | near-field | far-field | Provider-side speaker isolation — suppresses background voices and noise around the primary speaker before recognition. near-field suits handset calls, far-field suits speakerphone or distant mics. |
voiceFocusThreshold | number | 0.9 | 0–1 | How aggressively background audio is suppressed. Higher isolates harder but may clip quiet speech. Ignored when voiceFocus is off. |
Supported Languages
Language selection comes from the agent's ordered languages list (languages[0] = primary) — there is no per-provider language field. AssemblyAI supports these 12 platform languages:
| Code | Language | Code | Language | |
|---|---|---|---|---|
en | English | ru | Russian | |
zh | Chinese | de | German | |
es | Spanish | ja | Japanese | |
ar | Arabic | hi | Hindi | |
fr | French | mr | Marathi | |
pt | Portuguese | ur | Urdu |
An agent using AssemblyAI can select at most 10 of these languages — AssemblyAI's streaming API accepts up to 10 declared languages per session, so a config with more is rejected when it is saved.
Keyword boosting: AssemblyAI supports keyterms for domain-specific term boosting. Continuous partial transcripts are always on. Turn detection is fully provider-owned — the mode preset is the only turn-related control; there are no VAD or turn-silence dials.
Cartesia
Default model: ink-2
Models
| Model | Description |
|---|---|
ink-2 | Default -- English, Hindi, French, Japanese and Spanish, detected automatically (including Hinglish). Server-side turn detection. |
Configuration Fields
| Field | Type | Default | Range | Description |
|---|---|---|---|---|
model | string | ink-2 | -- | Recognition model. |
turnStartThreshold | float | 0.8 | 0.5 - 0.9 | Confidence at which a turn is considered started (Cartesia's own default is 0.8). |
turnEagerEndThreshold | float | 0.4 | 0.3 - 0.8 | Accepted but inactive: early turn end is not enabled, so this value does not affect turn timing (Cartesia's own default is 0.6). Not shown in the editor. |
turnEndThreshold | float | 0.2 | 0.05 - 0.5 | Confidence below which the turn ends (Cartesia's own default is 0.3). |
turnEndTimeoutMs | int | 2000 | 640 - 11200 | Force-ends the turn when the model stays uncertain — bounds the worst-case reply wait (Cartesia's own default is 5600). turnEndThreshold is the primary endpoint dial; if silence produces phantom turns, raise turnStartThreshold. |
keyterms | string | "" | -- | Comma-separated bias terms. |
Cartesia STT (ink-2) detects English, Hindi, French, Japanese and Spanish automatically and runs server-side turn detection with tunable thresholds. The thresholds must satisfy turnEndThreshold < turnEagerEndThreshold < turnStartThreshold.
Sarvam
Default model: saaras:v3
Models
| Model | Description |
|---|---|
saaras:v3 | Sarvam's recognition model |
Configuration Fields
| Field | Type | Default | Range | Description |
|---|---|---|---|---|
model | string | saaras:v3 | -- | Recognition model. |
mode | enum | transcribe | transcribe | verbatim | translit | codemix | Recognition mode. codemix handles Hindi/English code-mixed speech; translit transliterates to Latin script; verbatim keeps disfluencies. |
highVadSensitivity | bool | false | -- | Raise Sarvam's server VAD sensitivity: the turn ends after ~64 ms of silence instead of ~576 ms (16 kHz figures; roughly double on 8 kHz phone audio). |
Language selection comes from the agent's ordered languages list (languages[0] = primary) — there is no per-provider language field.
Supported Languages
en, hi, bn, gu, kn, ml, mr, or, pa, ta, te, ur, ne (Indian locales, xx-IN format).
Sarvam specializes in Indian languages. For real-world Indian voice agents that switch between Hindi and English mid-sentence (extremely common on phone calls), use mode=codemix — it is purpose-built for this. Sarvam runs on its server-side VAD defaults; highVadSensitivity is the only turn-taking dial. Pair Sarvam STT with Sarvam TTS (bulbul:v3) for the best end-to-end Indian language experience.