Agent Config schema
The complete nested agent configuration — one shape used by POST /agents, PATCH /agents, POST /call inline agents, and call snapshots.
AgentConfig is one schema used everywhere an agent is described:
POST /agentsandPATCH /agents/{uuid}— the saved agent'sconfigPOST /callagentfield — a transient inline agent (nothing saved)POST /agents/preview-snapshot— resolve and inspect without saving- Every call's stored
config_snapshot— the exact config the call ran with
There are no separate "saved" and "call-time" flavors: the same fields, defaults, and validation apply on every path. Field names are camelCase on the wire.
For the dashboard UI that produces this config, see Agents.
Minimal valid config
Only three things are required: a name and an llm block with provider + model. Everything else has a sensible default.
{
"name": "Support Agent",
"systemPrompt": "You help customers with orders. Be brief and warm.",
"llm": { "provider": "openai", "model": "gpt-5.4-mini" }
}name, llm, stt, and tts are required — for stt/tts, { "provider": ... } is enough (a null model resolves to the provider default). Omitted optional sections are filled with the defaults shown throughout this page, and an end_call tool is always seeded (see tools).
Unknown fields are rejected
The schema is strict (extra="forbid"): a misspelled or unknown key anywhere in the body fails with a 400 naming the exact path. Send only the fields defined on this page.
Top-level fields
| Field | Type | Default | Description |
|---|---|---|---|
name | string (1–120) | — required | Display name of the agent. |
description | string | "" | Free-form description (not spoken, not in the prompt). |
languages | string[] | ["en"] | Order is priority: languages[0] is the primary language, the rest are secondary hints. See Languages. |
systemPrompt | string | "" | LLM system prompt. Supports {{variable}} placeholders — see Context variables. |
timezone | string (IANA) | "Asia/Kolkata" | Always a concrete zone ("America/New_York", …). Drives the {{current_time}} / {{current_date}} built-ins. Invalid zones are rejected. |
greeting | object | see Greeting | Who speaks first and what the opener is. |
llm | object | — required | Model + the two sampling knobs. See LLM. |
stt | object | all defaults | Transcriber. See Transcriber (stt). |
tts | object | all defaults | Voice. See Voice (tts). |
noiseGate | object | enabled | Input cleanup before transcription. See Noise gate. |
backgroundSound | object | off | Ambient environment audio under the whole call. See Human-feel audio. |
taskSound | object | off | Work sounds while tools run. |
acknowledgeSounds | object | off | Spoken "hmm/okay" while the caller talks at length. |
fillerWords | object | off | Instant filler word over response latency. |
maxCallDurationS | int | 610 | 10..7200, multiple of 10. Hard call-length cap; the call ends with endedReason="max_duration". |
userIdleTimeoutS | int | 7 | 3..25 seconds of caller silence before a re-engage prompt. |
userIdleMessages | string[] (≤5) | ["Are you still there?"] | Escalating re-engage lines: retry N speaks entry N−1; retries past the end of the list speak the default line. Blank entries are dropped; an empty list resets to the default. |
userIdleMaxRetries | int | 2 | 1..5 re-engage attempts before hanging up (endedReason="silence_timeout"). |
tools | array | [end_call] | Inline tool objects and/or saved-tool UUID strings. See Tools. |
voicemailDetection | bool | true | Detect answering machines on outbound telephony calls. |
webhook | object | null | null | Lifecycle webhook — { "url": "https://…", "headers": { … } }. url must be http(s)://; a bare string is not accepted. |
enableSummarization | bool | false | Generate an AI summary at end of call, stored on the call row and included in the call.completed webhook. |
recordInProviderConsole | bool | null | null (= off) | Telephony only. Also record the full call on the carrier side (Twilio/Vobiz console). Zoxa's own recording is unaffected. |
contextVariables | object | {} | Saved defaults for the agent's {{variable}} placeholders — the bottom layer of call-time substitution. Max 200 keys, values ≤2000 chars. System built-ins (like user_number) are stripped — they can't be overridden. |
Languages
languages is an ordered list of language ids — the first entry is the primary language; there is no separate "primary" field. Later entries are secondary hints for code-switching (how they are consumed depends on the transcriber: Soniox, Flux multi, and AssemblyAI take the whole list as hints; Sarvam uses the single language, or auto-detects when several are selected; Cartesia detects the language itself).
Valid ids: en, hi, gu, ta, te, es, fr, ja, de, zh, ar, pt, ru, mr, bn, kn, ml, pa, ur, id, or, ne.
Validation is fail-closed: every listed language must be supported by both the chosen STT model and the chosen TTS model, or the request is rejected with a message naming the unsupported languages. Unknown ids are a hard error (never silently dropped — that could silently change your primary). Duplicates are deduplicated keeping the first occurrence. With AssemblyAI STT, at most 10 languages can be selected (AssemblyAI's per-session limit); more is rejected. Per-provider support matrices: STT providers · TTS providers.
{ "languages": ["en", "hi"] }Greeting
Controls the opening moments of every call.
{
"greeting": {
"firstMessages": ["Hi, this is Maya from Acme — how can I help?"],
"interruptible": false,
"speakFirst": "agent",
"agentDelayS": 0.0,
"userTimeoutS": 3.0
}
}| Field | Type | Default | Description |
|---|---|---|---|
firstMessages | string[] (≤5) | [] | Opener variants — one is chosen at random each call. Empty list = the LLM improvises the opener from the system prompt. |
interruptible | bool | false | Whether the caller can barge in over the opener. |
speakFirst | "agent" | "user" | "agent" | Who opens the call. |
agentDelayS | float | 0.0 | 0..4 s pause before the agent's opener (speakFirst: "agent"). |
userTimeoutS | float | 3.0 | 1..5 s to wait for the caller (speakFirst: "user") before the agent opens anyway. |
LLM
{
"llm": {
"provider": "openai",
"model": "gpt-5.4-mini",
"temperature": 1.0,
"maxTokens": 251,
"prewarm": true
}
}| Field | Type | Default | Description |
|---|---|---|---|
provider | string | — required | One of openai, qwen, google, anthropic. |
model | string | — required | Model id from the live catalog — GET /agents/models/available. Off-catalog ids are rejected. |
temperature | float | 1.0 | 0.0..2.0. Anthropic accepts only 0.0..1.0 — higher values are clamped to 1.0 on save. |
maxTokens | int | 251 | 1..4096. Reply-length rail: a spoken turn is a sentence or two, so the default stops runaway monologues without clipping real answers. |
prewarm | bool | true | Send one tiny hidden request at call start so the first real reply skips the provider's cold start. |
These are deliberately the only sampling knobs — topP, penalties, and seeds are not exposed; providers run on their own defaults.
Transcriber (stt)
Top-level selection plus one tuning block per provider. Only the block matching provider is used — the others may be present (handy when switching) and are simply ignored.
{
"stt": {
"provider": "soniox",
"model": null,
"interruptionMinWords": 0,
"soniox": { "contextTerms": "", "endpointSensitivity": 0.3, "endpointLatencyAdjustmentLevel": 2, "maxEndpointDelayMs": 2000 }
}
}| Field | Type | Default | Description |
|---|---|---|---|
provider | string | "soniox" | One of soniox, flux (Deepgram Flux), assemblyai, sarvam, cartesia. |
model | string | null | null | null = the provider's default model: Soniox stt-rt-v5, Flux flux-general-en (or flux-general-multi), AssemblyAI universal-3-6-pro, Sarvam saaras:v3, Cartesia ink-2. |
interruptionMinWords | int | 0 | 0..10. The only turn-taking knob. 0 = the caller's first detected speech interrupts the agent instantly. N≥1 = while the agent is speaking, the caller must get N transcribed words out before the agent stops — filters coughs and "yeah/ok" backchannels. |
Turn detection is provider-native
There is no turn-detection config. Each transcriber's own turn model decides when the caller has finished speaking — that's part of choosing a transcriber. The per-provider blocks below tune that provider's dials only.
stt.soniox
| Field | Default | Range | Description |
|---|---|---|---|
contextTerms | "" | — | Comma-separated bias terms (names, brands, jargon). |
endpointSensitivity | 0.3 | -1..1 | How readily the turn ends — higher is faster but may clip. |
endpointLatencyAdjustmentLevel | 2 | 0..3 | Trades a little accuracy for a faster turn end. |
maxEndpointDelayMs | 2000 | 500..3000 | Hard cap on post-pause wait before the turn ends. |
stt.flux
| Field | Default | Range | Description |
|---|---|---|---|
eotThreshold | 0.7 | 0.5..0.9 | Confidence needed to declare end of turn. |
eotTimeoutMs | 2000 | 500..60000 | Silence hard cap regardless of confidence. |
keyterms | "" | — | Comma-separated bias terms, hot-applied mid-stream. |
stt.assemblyai
| Field | Default | Range | Description |
|---|---|---|---|
mode | "balanced" | balanced | min_latency | max_accuracy | Preset that tunes the server's turn dynamics as one dial. |
keyterms | "" | — | Up to 100 comma-separated bias terms, ≤50 chars each. |
voiceFocus | "near-field" | off | near-field | far-field | Provider-side speaker isolation — suppresses background voices before recognition. near-field for handset calls, far-field for speakerphone. |
voiceFocusThreshold | 0.9 | 0..1 | Isolation aggressiveness. Ignored when voiceFocus is "off". |
stt.sarvam
| Field | Default | Range | Description |
|---|---|---|---|
mode | "transcribe" | transcribe | verbatim | translit | codemix | Rendering mode. When languages is exactly English + Hindi, codemix (Hinglish) is applied automatically. |
highVadSensitivity | false | — | Ends the turn after ~64 ms of silence instead of ~576 ms (16 kHz figures; roughly double on 8 kHz phone audio) — snappier, more prone to cutting the caller off. |
stt.cartesia
Thresholds must satisfy turnEndThreshold < turnEagerEndThreshold < turnStartThreshold — violations are rejected.
| Field | Default | Range | Description |
|---|---|---|---|
turnStartThreshold | 0.8 | 0.5..0.9 | Speech confidence needed to open the caller's turn (Cartesia's own default is 0.8). |
turnEagerEndThreshold | 0.4 | 0.3..0.8 | Accepted but inactive: early turn end is not enabled, so this value does not affect turn timing (Cartesia's own default is 0.6). Still counts toward the ordering rule above. |
turnEndThreshold | 0.2 | 0.05..0.5 | Confidence below which the turn ends — the primary endpoint dial (Cartesia's own default is 0.3). |
turnEndTimeoutMs | 2000 | 640..11200 | Force-ends the turn when the model stays uncertain. Lower = faster worst-case replies (Cartesia's own default is 5600). |
keyterms | "" | — | Up to 100 terms / 1200 chars total, comma-separated. |
Voice (tts)
Same pattern as the transcriber: top-level provider / model / voice, plus one tuning block per provider.
{
"tts": {
"provider": "elevenlabs",
"model": null,
"voice": null,
"elevenlabs": { "stability": 0.5, "similarityBoost": 0.75, "style": 0.0, "useSpeakerBoost": true, "speed": 1.0, "autoMode": true, "applyTextNormalization": "auto" }
}
}| Field | Type | Default | Description |
|---|---|---|---|
provider | string | "elevenlabs" | One of elevenlabs, cartesia, smallest, sarvam, xai, inworld. |
model | string | null | null | null = provider default: ElevenLabs eleven_v4_turbo, Cartesia sonic-3.6, Smallest lightning_v3.1, Sarvam bulbul:v3, Inworld inworld-tts-2. xAI has no model selector. |
voice | string | null | null | null = provider default voice. xAI accepts exactly ara, eve, leo, rex, sal — anything else is rejected. Other providers accept any id from their voice catalog. |
Per-provider voice tuning
| Block | Field | Default | Range |
|---|---|---|---|
tts.elevenlabs | stability | 0.5 | 0..1 (the only field eleven_v3_conversational and eleven_v4_turbo read) |
similarityBoost | 0.75 | 0..1 | |
style | 0.0 | 0..1 | |
useSpeakerBoost | true | — | |
speed | 1.0 | 0.7..1.2 | |
autoMode | true | — | |
applyTextNormalization | "auto" | auto | on | off | |
tts.cartesia | speed | 1.0 | 0.6..1.5 |
volume | 1.5 | 0.5..2.0 (1.0 leaves the voice unchanged) | |
tts.smallest | speed | 1.0 | 0.5..2.0 |
tts.sarvam | pace | 1.0 | 0.5..2.0 |
temperature | 0.6 | 0.01..1.0 | |
tts.xai | speed | 1.0 | 0.7..1.5 |
optimizeStreamingLatency | 0 | 0..2 | |
textNormalization | false | — | |
tts.inworld | speakingRate | 1.0 | 0.5..1.5 |
temperature | 1.0 | 0.01..2.0 (inworld-tts-2-flash only — Inworld ignores it on inworld-tts-2) |
Eleven v3 (eleven_v3_conversational) and Eleven v4 Turbo (eleven_v4_turbo) stream over ElevenLabs' Text-to-Dialogue connection, which accepts only stability. The other tts.elevenlabs fields are still validated and saved, but have no effect on those two models — they apply to eleven_flash_v2_5. To change speaking pace on v3 or v4 Turbo, use a different model; neither offers a speed control.
Noise gate
Input noise suppression / voice isolation applied to caller audio before transcription. On by default.
{ "noiseGate": { "enabled": true, "model": "voice-focus", "enhancementLevel": 0.8 } }| Field | Default | Description |
|---|---|---|
enabled | true | Turn the gate on/off. |
model | "voice-focus" | voice-focus isolates the nearest speaker and removes background voices; noise-suppression is general background-noise cleanup. |
enhancementLevel | 0.8 | 0..1 — processing strength. |
Human-feel audio
Four independent sections that make the agent sound less like a bot. All off by default — a disabled section has zero runtime footprint.
backgroundSound
Faint continuous ambience under the whole call — removes the tell-tale digital silence of a bot line.
| Field | Default | Range / Values |
|---|---|---|
enabled | false | — |
sound | "library" | crowd | library | restaurant | supermarket |
volume | 0.08 | 0.02..0.20 — above ~0.20 ambience reads as noise. |
taskSound
Keyboard typing or page turns while the agent runs tools — like a human agent working. Audio-only: what the agent says before a tool is owned by the tool's own messages. Starts after the tool's spoken message finishes (or straight into the silence for silent tools); stops the instant the agent speaks again.
| Field | Default | Range / Values |
|---|---|---|
enabled | false | — |
sound | "typing" | typing | page-turn |
volume | 0.35 | 0.05..0.60 |
startAfterMs | 800 | 0..5000 — fast tools finish before the sound ever plays. |
acknowledgeSounds
Short acknowledgements ("hmm", "okay") in the agent's own voice while the caller speaks at length. Never appears in transcripts.
| Field | Default | Range / Values |
|---|---|---|
enabled | false | — |
words | "hmm, hm hmm, okay" | Comma-separated; captured once per call in the agent's real voice. Required non-empty when enabled. |
wordOrder | "round-robin" | round-robin (in order; skipped chances don't advance) | random |
volume | 1.0 | 0.1..1.5 |
thresholdS | 3.0 | 2..15 — continuous caller speech per acknowledgement chance. |
probability | 0.8 | 0.05..1 — skipped chances keep it from sounding rhythmic. |
fillerWords
If the answer hasn't started within the threshold after the caller finishes, the agent instantly speaks a natural filler ("so", "yeah") to cover the thinking gap. The word does appear in the transcript — as the spoken prefix of the answer.
| Field | Default | Range / Values |
|---|---|---|
enabled | false | — |
words | "so, yeah, okay" | Comma-separated. Required non-empty when enabled. |
wordOrder | "round-robin" | round-robin | random |
thresholdMs | 300 | 100..3000 — response lag before a filler fires. |
probability | 0.75 | 0.05..1 |
There is deliberately no volume knob: the filler is the spoken prefix of the answer, so it always plays at exactly the voice's level.
Tools
tools is a single array mixing two kinds of entries:
- UUID strings — references to saved tools from
POST /tools - Inline tool objects — full definitions living inside the config
{
"tools": [
"b7e2a1c4-5d6f-4a8b-9c0d-1e2f3a4b5c6d",
{
"type": "function",
"name": "check_order_status",
"description": "Look up the status of the caller's order by order id.",
"config": {
"server": { "url": "https://api.acme.com/orders/status", "method": "POST" },
"parameters": { "type": "object", "properties": { "order_id": { "type": "string" } }, "required": ["order_id"] },
"messages": ["Let me pull that up for you."]
}
},
{
"type": "endCall",
"name": "end_call",
"description": "End the call when the user says goodbye or the conversation is finished.",
"config": { "messageType": "custom", "customMessages": ["Thanks for calling. Goodbye!"] }
}
]
}Inline types: function (HTTP), query (knowledge base), endCall, transferCall, testTool. Full field reference for each: Tool types.
Rules enforced at validation:
endCallis mandatory and auto-seeded. If your config has noendCalltool, the default one (name: "end_call", farewell"Goodbye!") is appended automatically — every stored and returned config includes it.- Tool names are 1–64 chars of
[a-zA-Z0-9_-], must be unique after sanitization to LLM-function form (book-slotandbook_slotcollide), and must not use the reserved namesend_call/retrieve_from_knowledge_base/safe_calculator(exceptendCall-type tools, which may useend_call). - Spoken
messagesarrays on tools follow the same variant rules as everywhere else: ≤5 entries, blanks dropped, ONE spoken per invocation. Tool messages additionally carry amessagesOrderknob ("random"default /"in-order"= sequential within a call, wrapping); greeting and endCall variants are always random.
Message variants
Every spoken static message in the config is a list of up to 5 entries. How the list is consumed is fixed per context:
| Field | Consumption |
|---|---|
greeting.firstMessages | Random pick per call; [] = LLM improvises. |
endCall customMessages | Random pick when the farewell is spoken. |
function / query tool messages | Sequential per invocation, wraps around; [] = silent. |
userIdleMessages | Sequential per retry; retries past the end speak "Are you still there?". |
Validation summary
- Strict schema — unknown keys anywhere fail with
400and the exact field path. languages: unknown ids rejected; every language must be supported by both the chosen STT and TTS model; at most 10 with AssemblyAI STT.llm.temperatureabove1.0is clamped for Anthropic.- Cartesia STT threshold ordering (
end < eagerEnd < start) is enforced. - xAI voices and Inworld model ids are validated against their fixed catalogs.
- HTTP tool URLs must be
http(s)://and may not target private/internal IPs or metadata endpoints. endCallis seeded when absent;messageType: "custom"requires a non-emptycustomMessages,"audio"requiresaudioRecordingId.
Preview before saving
Use POST /agents/preview-snapshot to see the fully-resolved config — all defaults materialized, endCall seeded — without saving anything. What it returns is exactly what a call would snapshot.