zoxaAI
Homepage

STT Providers

Every speech-to-text provider supported by zoxaAI -- models, language support, and configuration options per provider.

Overview

zoxaAI supports 5 STT providers (6 models). Each agent is configured with exactly one STT provider. Language support, model selection, and recognition features vary by provider. Models are a curated catalog — you pick from the models listed per provider below.


Provider Comparison

ProviderDefault ModelLanguagesKeyword BoostingNotes
Sonioxstt-rt-v556+Yes (contextTerms)Platform default. Native code-switching — the agent's whole languages list stays live in one stream, no model swap
Deepgram Fluxflux-general-enEnglish (flux-general-en) / 8 (flux-general-multi)Yes (keyterms)Server-side end-of-turn detection
AssemblyAIuniversal-3-6-pro12 (of the model's 32)Yes (keyterms)High accuracy; latency/accuracy mode presets
Cartesiaink-25 (auto-detected)Yes (keyterms)Server-side turn detection with tunable thresholds
Sarvamsaaras:v313 (Indian locales) + code-mixNoIndian language specialist; server-side VAD

Soniox

Default model: stt-rt-v5 — the platform default STT provider.

Models

ModelNotes
stt-rt-v5Default — single realtime streaming model (Soniox's stt-rt-v5).

Configuration Fields

FieldTypeDefaultRangeDescription
modelstringstt-rt-v5--Recognition model.
contextTermsstring""--Comma-separated bias terms (names, jargon) to improve recognition. Empty = off.
endpointSensitivityfloat0.3-1.0 - 1.0How eagerly Soniox ends a turn. Higher = faster endpointing.
endpointLatencyAdjustmentLevelint20 - 3Latency-vs-accuracy tradeoff for endpoint detection.
maxEndpointDelayMsint2000500 - 3000Upper bound on how long Soniox waits before forcing a turn end (platform-uniform 2s silence cap).

Supported Languages

Soniox recognizes 56+ languages in a single realtime model. Language selection comes from the agent's ordered languages list — Soniox receives the whole list as live hints, including en, zh, es, ar, fr, pt, ru, id, de, ja, hi, bn, gu, kn, ml, mr, pa, ta, te, ur.

Why Soniox for multilingual phone calls

Most STTs that claim "multilingual" pick one language up front and recognize that one. Soniox keeps every language in the agent's languages list live in the same stream — the model expects code-switching mid-utterance and handles it natively. That makes it the best fit for Indian voice agents (English ↔ Hindi mid-sentence) and any market where speakers routinely mix two languages.

Example

{
  "languages": ["en", "hi"],
  "stt": {
    "provider": "soniox",
    "soniox": { "contextTerms": "Acme, zoxa" }
  }
}

Deepgram Flux

Default model: flux-general-en

Models

ModelDescription
flux-general-enDefault -- English only. Hides the language picker.
flux-general-multiMultilingual (en, es, fr, pt, ru, de, ja, hi).

Configuration Fields

FieldTypeDefaultRangeDescription
modelstringflux-general-en--Recognition model.
eotThresholdfloat0.70.5 - 0.9End-of-turn confidence needed before the turn is finalized.
eotTimeoutMsint2000500 - 60000Max wait before the turn is force-ended (platform-uniform 2s silence cap).
keytermsstring""--Comma-separated bias terms, hot-applied mid-stream by Flux.

Flux runs server-side end-of-turn detection instead of a local silence timer. flux-general-en only supports English and hides the language picker; switch to flux-general-multi for the multilingual set. Both models support keyword boosting via keyterms.


AssemblyAI

Default model: universal-3-6-pro

Models

ModelDescription
universal-3-6-proDefault -- Universal-3.6 Pro Realtime: streams 32 languages with native code-switching; 12 of them are platform languages (below).

Configuration Fields

FieldTypeDefaultRangeDescription
modelstringuniversal-3-6-pro--Recognition model.
modeenumbalancedbalanced | min_latency | max_accuracyPreset that shifts several server defaults at once — min_latency responds fastest, max_accuracy waits longest before ending a turn.
keytermsstring""--Comma-separated bias terms (up to 100 terms / 50 chars each).
voiceFocusenumnear-fieldoff | near-field | far-fieldProvider-side speaker isolation — suppresses background voices and noise around the primary speaker before recognition. near-field suits handset calls, far-field suits speakerphone or distant mics.
voiceFocusThresholdnumber0.90–1How aggressively background audio is suppressed. Higher isolates harder but may clip quiet speech. Ignored when voiceFocus is off.

Supported Languages

Language selection comes from the agent's ordered languages list (languages[0] = primary) — there is no per-provider language field. AssemblyAI supports these 12 platform languages:

CodeLanguageCodeLanguage
enEnglishruRussian
zhChinesedeGerman
esSpanishjaJapanese
arArabichiHindi
frFrenchmrMarathi
ptPortugueseurUrdu

An agent using AssemblyAI can select at most 10 of these languages — AssemblyAI's streaming API accepts up to 10 declared languages per session, so a config with more is rejected when it is saved.

Keyword boosting: AssemblyAI supports keyterms for domain-specific term boosting. Continuous partial transcripts are always on. Turn detection is fully provider-owned — the mode preset is the only turn-related control; there are no VAD or turn-silence dials.


Cartesia

Default model: ink-2

Models

ModelDescription
ink-2Default -- English, Hindi, French, Japanese and Spanish, detected automatically (including Hinglish). Server-side turn detection.

Configuration Fields

FieldTypeDefaultRangeDescription
modelstringink-2--Recognition model.
turnStartThresholdfloat0.80.5 - 0.9Confidence at which a turn is considered started (Cartesia's own default is 0.8).
turnEagerEndThresholdfloat0.40.3 - 0.8Accepted but inactive: early turn end is not enabled, so this value does not affect turn timing (Cartesia's own default is 0.6). Not shown in the editor.
turnEndThresholdfloat0.20.05 - 0.5Confidence below which the turn ends (Cartesia's own default is 0.3).
turnEndTimeoutMsint2000640 - 11200Force-ends the turn when the model stays uncertain — bounds the worst-case reply wait (Cartesia's own default is 5600). turnEndThreshold is the primary endpoint dial; if silence produces phantom turns, raise turnStartThreshold.
keytermsstring""--Comma-separated bias terms.

Cartesia STT (ink-2) detects English, Hindi, French, Japanese and Spanish automatically and runs server-side turn detection with tunable thresholds. The thresholds must satisfy turnEndThreshold < turnEagerEndThreshold < turnStartThreshold.


Sarvam

Default model: saaras:v3

Models

ModelDescription
saaras:v3Sarvam's recognition model

Configuration Fields

FieldTypeDefaultRangeDescription
modelstringsaaras:v3--Recognition model.
modeenumtranscribetranscribe | verbatim | translit | codemixRecognition mode. codemix handles Hindi/English code-mixed speech; translit transliterates to Latin script; verbatim keeps disfluencies.
highVadSensitivityboolfalse--Raise Sarvam's server VAD sensitivity: the turn ends after ~64 ms of silence instead of ~576 ms (16 kHz figures; roughly double on 8 kHz phone audio).

Language selection comes from the agent's ordered languages list (languages[0] = primary) — there is no per-provider language field.

Supported Languages

en, hi, bn, gu, kn, ml, mr, or, pa, ta, te, ur, ne (Indian locales, xx-IN format).

Sarvam specializes in Indian languages. For real-world Indian voice agents that switch between Hindi and English mid-sentence (extremely common on phone calls), use mode=codemix — it is purpose-built for this. Sarvam runs on its server-side VAD defaults; highVadSensitivity is the only turn-taking dial. Pair Sarvam STT with Sarvam TTS (bulbul:v3) for the best end-to-end Indian language experience.


Choosing a Provider

On this page