zoxaAI
Homepage
API ReferenceSchemas

Agent Config schema

The complete nested agent configuration — one shape used by POST /agents, PATCH /agents, POST /call inline agents, and call snapshots.

AgentConfig is one schema used everywhere an agent is described:

There are no separate "saved" and "call-time" flavors: the same fields, defaults, and validation apply on every path. Field names are camelCase on the wire.

For the dashboard UI that produces this config, see Agents.

Minimal valid config

Only three things are required: a name and an llm block with provider + model. Everything else has a sensible default.

{
  "name": "Support Agent",
  "systemPrompt": "You help customers with orders. Be brief and warm.",
  "llm": { "provider": "openai", "model": "gpt-5.4-mini" }
}

name, llm, stt, and tts are required — for stt/tts, { "provider": ... } is enough (a null model resolves to the provider default). Omitted optional sections are filled with the defaults shown throughout this page, and an end_call tool is always seeded (see tools).

Unknown fields are rejected

The schema is strict (extra="forbid"): a misspelled or unknown key anywhere in the body fails with a 400 naming the exact path. Send only the fields defined on this page.

Top-level fields

FieldTypeDefaultDescription
namestring (1–120)— requiredDisplay name of the agent.
descriptionstring""Free-form description (not spoken, not in the prompt).
languagesstring[]["en"]Order is priority: languages[0] is the primary language, the rest are secondary hints. See Languages.
systemPromptstring""LLM system prompt. Supports {{variable}} placeholders — see Context variables.
timezonestring (IANA)"Asia/Kolkata"Always a concrete zone ("America/New_York", …). Drives the {{current_time}} / {{current_date}} built-ins. Invalid zones are rejected.
greetingobjectsee GreetingWho speaks first and what the opener is.
llmobject— requiredModel + the two sampling knobs. See LLM.
sttobjectall defaultsTranscriber. See Transcriber (stt).
ttsobjectall defaultsVoice. See Voice (tts).
noiseGateobjectenabledInput cleanup before transcription. See Noise gate.
backgroundSoundobjectoffAmbient environment audio under the whole call. See Human-feel audio.
taskSoundobjectoffWork sounds while tools run.
acknowledgeSoundsobjectoffSpoken "hmm/okay" while the caller talks at length.
fillerWordsobjectoffInstant filler word over response latency.
maxCallDurationSint61010..7200, multiple of 10. Hard call-length cap; the call ends with endedReason="max_duration".
userIdleTimeoutSint73..25 seconds of caller silence before a re-engage prompt.
userIdleMessagesstring[] (≤5)["Are you still there?"]Escalating re-engage lines: retry N speaks entry N−1; retries past the end of the list speak the default line. Blank entries are dropped; an empty list resets to the default.
userIdleMaxRetriesint21..5 re-engage attempts before hanging up (endedReason="silence_timeout").
toolsarray[end_call]Inline tool objects and/or saved-tool UUID strings. See Tools.
voicemailDetectionbooltrueDetect answering machines on outbound telephony calls.
webhookobject | nullnullLifecycle webhook — { "url": "https://…", "headers": { … } }. url must be http(s)://; a bare string is not accepted.
enableSummarizationboolfalseGenerate an AI summary at end of call, stored on the call row and included in the call.completed webhook.
recordInProviderConsolebool | nullnull (= off)Telephony only. Also record the full call on the carrier side (Twilio/Vobiz console). Zoxa's own recording is unaffected.
contextVariablesobject{}Saved defaults for the agent's {{variable}} placeholders — the bottom layer of call-time substitution. Max 200 keys, values ≤2000 chars. System built-ins (like user_number) are stripped — they can't be overridden.

Languages

languages is an ordered list of language ids — the first entry is the primary language; there is no separate "primary" field. Later entries are secondary hints for code-switching (how they are consumed depends on the transcriber: Soniox, Flux multi, and AssemblyAI take the whole list as hints; Sarvam uses the single language, or auto-detects when several are selected; Cartesia detects the language itself).

Valid ids: en, hi, gu, ta, te, es, fr, ja, de, zh, ar, pt, ru, mr, bn, kn, ml, pa, ur, id, or, ne.

Validation is fail-closed: every listed language must be supported by both the chosen STT model and the chosen TTS model, or the request is rejected with a message naming the unsupported languages. Unknown ids are a hard error (never silently dropped — that could silently change your primary). Duplicates are deduplicated keeping the first occurrence. With AssemblyAI STT, at most 10 languages can be selected (AssemblyAI's per-session limit); more is rejected. Per-provider support matrices: STT providers · TTS providers.

{ "languages": ["en", "hi"] }

Greeting

Controls the opening moments of every call.

{
  "greeting": {
    "firstMessages": ["Hi, this is Maya from Acme — how can I help?"],
    "interruptible": false,
    "speakFirst": "agent",
    "agentDelayS": 0.0,
    "userTimeoutS": 3.0
  }
}
FieldTypeDefaultDescription
firstMessagesstring[] (≤5)[]Opener variants — one is chosen at random each call. Empty list = the LLM improvises the opener from the system prompt.
interruptibleboolfalseWhether the caller can barge in over the opener.
speakFirst"agent" | "user""agent"Who opens the call.
agentDelaySfloat0.00..4 s pause before the agent's opener (speakFirst: "agent").
userTimeoutSfloat3.01..5 s to wait for the caller (speakFirst: "user") before the agent opens anyway.

LLM

{
  "llm": {
    "provider": "openai",
    "model": "gpt-5.4-mini",
    "temperature": 1.0,
    "maxTokens": 251,
    "prewarm": true
  }
}
FieldTypeDefaultDescription
providerstring— requiredOne of openai, qwen, google, anthropic.
modelstring— requiredModel id from the live catalog — GET /agents/models/available. Off-catalog ids are rejected.
temperaturefloat1.00.0..2.0. Anthropic accepts only 0.0..1.0 — higher values are clamped to 1.0 on save.
maxTokensint2511..4096. Reply-length rail: a spoken turn is a sentence or two, so the default stops runaway monologues without clipping real answers.
prewarmbooltrueSend one tiny hidden request at call start so the first real reply skips the provider's cold start.

These are deliberately the only sampling knobs — topP, penalties, and seeds are not exposed; providers run on their own defaults.

Transcriber (stt)

Top-level selection plus one tuning block per provider. Only the block matching provider is used — the others may be present (handy when switching) and are simply ignored.

{
  "stt": {
    "provider": "soniox",
    "model": null,
    "interruptionMinWords": 0,
    "soniox": { "contextTerms": "", "endpointSensitivity": 0.3, "endpointLatencyAdjustmentLevel": 2, "maxEndpointDelayMs": 2000 }
  }
}
FieldTypeDefaultDescription
providerstring"soniox"One of soniox, flux (Deepgram Flux), assemblyai, sarvam, cartesia.
modelstring | nullnullnull = the provider's default model: Soniox stt-rt-v5, Flux flux-general-en (or flux-general-multi), AssemblyAI universal-3-6-pro, Sarvam saaras:v3, Cartesia ink-2.
interruptionMinWordsint00..10. The only turn-taking knob. 0 = the caller's first detected speech interrupts the agent instantly. N≥1 = while the agent is speaking, the caller must get N transcribed words out before the agent stops — filters coughs and "yeah/ok" backchannels.

Turn detection is provider-native

There is no turn-detection config. Each transcriber's own turn model decides when the caller has finished speaking — that's part of choosing a transcriber. The per-provider blocks below tune that provider's dials only.

stt.soniox

FieldDefaultRangeDescription
contextTerms""—Comma-separated bias terms (names, brands, jargon).
endpointSensitivity0.3-1..1How readily the turn ends — higher is faster but may clip.
endpointLatencyAdjustmentLevel20..3Trades a little accuracy for a faster turn end.
maxEndpointDelayMs2000500..3000Hard cap on post-pause wait before the turn ends.

stt.flux

FieldDefaultRangeDescription
eotThreshold0.70.5..0.9Confidence needed to declare end of turn.
eotTimeoutMs2000500..60000Silence hard cap regardless of confidence.
keyterms""—Comma-separated bias terms, hot-applied mid-stream.

stt.assemblyai

FieldDefaultRangeDescription
mode"balanced"balanced | min_latency | max_accuracyPreset that tunes the server's turn dynamics as one dial.
keyterms""—Up to 100 comma-separated bias terms, ≤50 chars each.
voiceFocus"near-field"off | near-field | far-fieldProvider-side speaker isolation — suppresses background voices before recognition. near-field for handset calls, far-field for speakerphone.
voiceFocusThreshold0.90..1Isolation aggressiveness. Ignored when voiceFocus is "off".

stt.sarvam

FieldDefaultRangeDescription
mode"transcribe"transcribe | verbatim | translit | codemixRendering mode. When languages is exactly English + Hindi, codemix (Hinglish) is applied automatically.
highVadSensitivityfalse—Ends the turn after ~64 ms of silence instead of ~576 ms (16 kHz figures; roughly double on 8 kHz phone audio) — snappier, more prone to cutting the caller off.

stt.cartesia

Thresholds must satisfy turnEndThreshold < turnEagerEndThreshold < turnStartThreshold — violations are rejected.

FieldDefaultRangeDescription
turnStartThreshold0.80.5..0.9Speech confidence needed to open the caller's turn (Cartesia's own default is 0.8).
turnEagerEndThreshold0.40.3..0.8Accepted but inactive: early turn end is not enabled, so this value does not affect turn timing (Cartesia's own default is 0.6). Still counts toward the ordering rule above.
turnEndThreshold0.20.05..0.5Confidence below which the turn ends — the primary endpoint dial (Cartesia's own default is 0.3).
turnEndTimeoutMs2000640..11200Force-ends the turn when the model stays uncertain. Lower = faster worst-case replies (Cartesia's own default is 5600).
keyterms""—Up to 100 terms / 1200 chars total, comma-separated.

Voice (tts)

Same pattern as the transcriber: top-level provider / model / voice, plus one tuning block per provider.

{
  "tts": {
    "provider": "elevenlabs",
    "model": null,
    "voice": null,
    "elevenlabs": { "stability": 0.5, "similarityBoost": 0.75, "style": 0.0, "useSpeakerBoost": true, "speed": 1.0, "autoMode": true, "applyTextNormalization": "auto" }
  }
}
FieldTypeDefaultDescription
providerstring"elevenlabs"One of elevenlabs, cartesia, smallest, sarvam, xai, inworld.
modelstring | nullnullnull = provider default: ElevenLabs eleven_v4_turbo, Cartesia sonic-3.6, Smallest lightning_v3.1, Sarvam bulbul:v3, Inworld inworld-tts-2. xAI has no model selector.
voicestring | nullnullnull = provider default voice. xAI accepts exactly ara, eve, leo, rex, sal — anything else is rejected. Other providers accept any id from their voice catalog.

Per-provider voice tuning

BlockFieldDefaultRange
tts.elevenlabsstability0.50..1 (the only field eleven_v3_conversational and eleven_v4_turbo read)
similarityBoost0.750..1
style0.00..1
useSpeakerBoosttrue—
speed1.00.7..1.2
autoModetrue—
applyTextNormalization"auto"auto | on | off
tts.cartesiaspeed1.00.6..1.5
volume1.50.5..2.0 (1.0 leaves the voice unchanged)
tts.smallestspeed1.00.5..2.0
tts.sarvampace1.00.5..2.0
temperature0.60.01..1.0
tts.xaispeed1.00.7..1.5
optimizeStreamingLatency00..2
textNormalizationfalse—
tts.inworldspeakingRate1.00.5..1.5
temperature1.00.01..2.0 (inworld-tts-2-flash only — Inworld ignores it on inworld-tts-2)

Eleven v3 (eleven_v3_conversational) and Eleven v4 Turbo (eleven_v4_turbo) stream over ElevenLabs' Text-to-Dialogue connection, which accepts only stability. The other tts.elevenlabs fields are still validated and saved, but have no effect on those two models — they apply to eleven_flash_v2_5. To change speaking pace on v3 or v4 Turbo, use a different model; neither offers a speed control.

Noise gate

Input noise suppression / voice isolation applied to caller audio before transcription. On by default.

{ "noiseGate": { "enabled": true, "model": "voice-focus", "enhancementLevel": 0.8 } }
FieldDefaultDescription
enabledtrueTurn the gate on/off.
model"voice-focus"voice-focus isolates the nearest speaker and removes background voices; noise-suppression is general background-noise cleanup.
enhancementLevel0.80..1 — processing strength.

Human-feel audio

Four independent sections that make the agent sound less like a bot. All off by default — a disabled section has zero runtime footprint.

backgroundSound

Faint continuous ambience under the whole call — removes the tell-tale digital silence of a bot line.

FieldDefaultRange / Values
enabledfalse—
sound"library"crowd | library | restaurant | supermarket
volume0.080.02..0.20 — above ~0.20 ambience reads as noise.

taskSound

Keyboard typing or page turns while the agent runs tools — like a human agent working. Audio-only: what the agent says before a tool is owned by the tool's own messages. Starts after the tool's spoken message finishes (or straight into the silence for silent tools); stops the instant the agent speaks again.

FieldDefaultRange / Values
enabledfalse—
sound"typing"typing | page-turn
volume0.350.05..0.60
startAfterMs8000..5000 — fast tools finish before the sound ever plays.

acknowledgeSounds

Short acknowledgements ("hmm", "okay") in the agent's own voice while the caller speaks at length. Never appears in transcripts.

FieldDefaultRange / Values
enabledfalse—
words"hmm, hm hmm, okay"Comma-separated; captured once per call in the agent's real voice. Required non-empty when enabled.
wordOrder"round-robin"round-robin (in order; skipped chances don't advance) | random
volume1.00.1..1.5
thresholdS3.02..15 — continuous caller speech per acknowledgement chance.
probability0.80.05..1 — skipped chances keep it from sounding rhythmic.

fillerWords

If the answer hasn't started within the threshold after the caller finishes, the agent instantly speaks a natural filler ("so", "yeah") to cover the thinking gap. The word does appear in the transcript — as the spoken prefix of the answer.

FieldDefaultRange / Values
enabledfalse—
words"so, yeah, okay"Comma-separated. Required non-empty when enabled.
wordOrder"round-robin"round-robin | random
thresholdMs300100..3000 — response lag before a filler fires.
probability0.750.05..1

There is deliberately no volume knob: the filler is the spoken prefix of the answer, so it always plays at exactly the voice's level.

Tools

tools is a single array mixing two kinds of entries:

  • UUID strings — references to saved tools from POST /tools
  • Inline tool objects — full definitions living inside the config
{
  "tools": [
    "b7e2a1c4-5d6f-4a8b-9c0d-1e2f3a4b5c6d",
    {
      "type": "function",
      "name": "check_order_status",
      "description": "Look up the status of the caller's order by order id.",
      "config": {
        "server": { "url": "https://api.acme.com/orders/status", "method": "POST" },
        "parameters": { "type": "object", "properties": { "order_id": { "type": "string" } }, "required": ["order_id"] },
        "messages": ["Let me pull that up for you."]
      }
    },
    {
      "type": "endCall",
      "name": "end_call",
      "description": "End the call when the user says goodbye or the conversation is finished.",
      "config": { "messageType": "custom", "customMessages": ["Thanks for calling. Goodbye!"] }
    }
  ]
}

Inline types: function (HTTP), query (knowledge base), endCall, transferCall, testTool. Full field reference for each: Tool types.

Rules enforced at validation:

  • endCall is mandatory and auto-seeded. If your config has no endCall tool, the default one (name: "end_call", farewell "Goodbye!") is appended automatically — every stored and returned config includes it.
  • Tool names are 1–64 chars of [a-zA-Z0-9_-], must be unique after sanitization to LLM-function form (book-slot and book_slot collide), and must not use the reserved names end_call / retrieve_from_knowledge_base / safe_calculator (except endCall-type tools, which may use end_call).
  • Spoken messages arrays on tools follow the same variant rules as everywhere else: ≤5 entries, blanks dropped, ONE spoken per invocation. Tool messages additionally carry a messagesOrder knob ("random" default / "in-order" = sequential within a call, wrapping); greeting and endCall variants are always random.

Message variants

Every spoken static message in the config is a list of up to 5 entries. How the list is consumed is fixed per context:

FieldConsumption
greeting.firstMessagesRandom pick per call; [] = LLM improvises.
endCall customMessagesRandom pick when the farewell is spoken.
function / query tool messagesSequential per invocation, wraps around; [] = silent.
userIdleMessagesSequential per retry; retries past the end speak "Are you still there?".

Validation summary

  • Strict schema — unknown keys anywhere fail with 400 and the exact field path.
  • languages: unknown ids rejected; every language must be supported by both the chosen STT and TTS model; at most 10 with AssemblyAI STT.
  • llm.temperature above 1.0 is clamped for Anthropic.
  • Cartesia STT threshold ordering (end < eagerEnd < start) is enforced.
  • xAI voices and Inworld model ids are validated against their fixed catalogs.
  • HTTP tool URLs must be http(s):// and may not target private/internal IPs or metadata endpoints.
  • endCall is seeded when absent; messageType: "custom" requires a non-empty customMessages, "audio" requires audioRecordingId.

Preview before saving

Use POST /agents/preview-snapshot to see the fully-resolved config — all defaults materialized, endCall seeded — without saving anything. What it returns is exactly what a call would snapshot.

On this page