zoxaAI
Homepage

Agents

Create and configure standalone voice AI agents with full control over LLM, voice, transcription, and call behavior.

What you'll learn

  • How to configure every setting in the agent editor across its five tabs
  • How the right-panel chips pick your transcriber, model, and voice — with live cost and latency stats
  • What each field does, its type, default value, and valid range
  • How to test your agent with browser calls, telephony simulation, phone calls, and the code preview

Agents are standalone voice AI entities that handle conversations end-to-end. Each agent has its own personality (system prompt), voice, language model, and call behavior settings. You configure an agent once, attach a phone number or talk to it in the browser, and it handles calls autonomously.


Agent Editor

The agent editor has two halves. The middle panel is organized into five tabs — Agent, Engine, Tools, Call, and Analytics — that group related configuration. The right panel is a live test surface: three clickable chips select your transcriber, language model, and voice; a cost/latency estimate updates as you pick; and a talk panel runs a real browser call against the agent you are editing.

Where model selection lives

You don't pick your provider, model, or voice on a tab — you pick them from the three chips in the right panel (Transcriber · Model · Voice). Every setting that is specific to the chosen model then lives in the Engine tab. The middle tabs never select a model; they configure behavior around whatever the chips have selected.


Right Panel: Transcriber, Model & Voice

The right panel shows three chips — Transcriber, Model, and Voice — reflecting the agent's current speech-to-text, LLM, and text-to-speech selection. Click any chip to expand an inline picker beneath it.

ChipSelectsConfig path
TranscriberSTT provider + modelstt.provider, stt.model
ModelLLM provider + modelllm.provider, llm.model
VoiceTTS provider + model + voicetts.provider, tts.model, tts.voice

Each picker lists the available providers (the curated lab set), then the models under the selected provider with per-model stats — the same cost and latency data that drives the Cost and Latency bars in the estimate panel. The Voice picker adds a voice list for the chosen TTS provider. Selecting a value updates the estimate bars live.

Cascade resets. Changing a provider resets the dependent fields to that provider's defaults: pick a new STT/LLM provider and the model resets to the provider default; pick a new TTS provider and both the model and voice reset. This guarantees you never save a model or voice that doesn't belong to the selected provider.

Coverage gating. Models that cannot cover your selected languages are greyed out in these pickers with a tooltip explaining why (see Languages below). Only one picker is open at a time; the chip shows an active state while open and the picker closes on select or on re-click.

The provider and model catalog is served by GET /api/v1/user/configurations/providers and is env-gated server-side, so every option the chips show is guaranteed usable on the platform. For the full per-provider capability tables, see STT Providers, TTS Providers, and LLM Providers.


Agent Tab

The Agent tab contains the core personality and behavior settings -- languages, the welcome message, the system prompt, and the End Call behavior.

Languages

An agent speaks an ordered list of languages — the order is the priority: the first language is the primary, the second is the secondary, and so on.

FieldTypeDefaultDescription
Languagesstring[]["en"]The languages the agent handles, ordered by priority — languages[0] is the primary. In the editor the primary chip wears a filled star; clicking a non-primary chip's star moves it to the front. The primary cannot be removed — star another language first. New languages are added at the end (lowest priority). Config path: config.languages.

The chosen list drives coverage gating in the right-panel Transcriber and Voice pickers: any STT or TTS model that cannot cover every selected language is greyed out with a tooltip. For example, picking hi + en greys out English-only models. Pick languages first, then choose models that cover them. The API enforces the same rule — saving or calling with a language the chosen transcriber or voice model does not support returns a 400 naming the unsupported languages.

Two-language calls

If your callers routinely mix two languages mid-sentence (Hinglish, Spanglish, etc.), add both languages and pick a transcriber that covers the pair. Soniox is the strongest fit — it takes every selected language as a live hint and handles mid-utterance code-switching in a single stream. See STT Providers → Soniox.

Welcome Message

The first thing your agent says when a call starts.

FieldTypeDefaultRangeDescription
Messagestringempty--The greeting spoken at the start of a call. Leave empty to let the agent generate a greeting from the system prompt. On the API this is firstMessages — an array of up to 5 variants; one is picked at random on each call, so repeat callers don't hear the identical opener. The editor binds the first entry; extra variants are API-only.
Interruptiblebooleanfalse--Whether the user can interrupt the welcome message while it is being spoken.
Who Speaks First"agent" or "user""agent"--Controls whether the agent speaks first or waits for the user.
Start After Delayfloat00 -- 4.0, step 0.1Visible when Agent speaks first is selected. Delays the agent's opening message by this many seconds. Useful for telephony scenarios where a short pause feels more natural.
Silence Timeoutfloat3.01.0 -- 5.0, step 0.5Visible when User speaks first is selected. If the user does not speak within this time, the agent starts with the welcome message anyway.

Tip: If you are building a customer support agent that handles inbound calls, set "Who Speaks First" to Agent with a short, friendly greeting. For outbound calls where you need to confirm you have the right person, set it to User so the agent waits for a "hello" before introducing itself.

System Prompt

The system prompt defines your agent's role, personality, knowledge, and behavior. This is the primary instruction set that shapes every response.

FieldTypeMax LengthDescription
Promptstring100,000 charsThe core instructions for the agent. Supports multi-line text.

The editor includes an AI Generate button that can automatically draft a system prompt based on your agent's name and description.

The Agent tab also carries a timezone picker (used to resolve system-time variables in the prompt) and the pill-based {{variable}} editor for the system prompt. The timezone is always a concrete IANA zone: new agents are created with your browser's timezone, and API bodies that omit timezone default to Asia/Kolkata.

Variables

Write {{variable_name}} anywhere in the system prompt, the welcome message, or the end-call trigger prompt, and the platform replaces it with a real value when the call starts. Matching is case- and space-insensitive: {{First Name}}, {{first_name}} and {{FIRST_NAME}} all resolve to the same value.

There are two kinds.

System variables

Resolved by the platform itself on every call. You never set these — type {{ in any prompt field to pick them from the menu.

VariableResolves toExample
{{current_time}}Current time in the agent's timezone, 24h12:33
{{current_day}}Current weekdayFriday
{{current_date}}Current date26 July 2026
{{current_timezone}}The agent's configured timezoneAsia/Kolkata
{{user_number}}The caller's phone number+14155550100
{{agent_number}}The agent's phone number+18885551234

System variables cannot be overridden

These six names have top-level authority. A saved default, or a contextVariables value sent by a campaign or the API, is ignored for these keys — the platform's own value always wins. This is deliberate: it guarantees {{current_time}} really is the current time and {{user_number}} really is the caller, on every single call.

{{user_number}} and {{agent_number}} follow the human, not the dialer — on an outbound call user_number is the number you dialed; on inbound it is the number that called you.

Your own variables (context variables)

Every other {{token}} is yours. Open Test Inputs (the {} button in the right panel) to see every variable found in your prompts, give each one a value, then save the agent — the values persist in the agent config as contextVariables and are reused on every future call. Reopen the agent and the rows come back filled in.

Saved values are defaults, not constants: a campaign row or an API call can override any of them per call.

Resolution order

Lowest to highest — the highest-priority source that produces a value wins:

PrioritySourceRule
1Nothing declares the nameThe literal {{token}} is left in the prompt, so a typo stays visible instead of silently vanishing
2Saved contextVariables on the agentUsed as the default. A saved blank is honored — it renders as empty text rather than leaving {{token}} in the prompt
3Campaign row / API contextVariablesOverrides the saved default, but only when non-blank. A blank CSV cell or empty string counts as "not provided" and falls back to the saved default
4System variableAlways wins — see the warning above

Why blank counts as not provided

Campaign CSVs are often sparse. If one row is missing a city value you almost never want that call to say "calling from ." — you want it to fall back to whatever you saved on the agent. So an empty value from a campaign or API request is treated as absent, and the saved default is used instead.

If you genuinely want a variable to render as nothing, save it as blank on the agent (priority 2).

For example, with the prompt Calling about your {{plan}} plan in {{city}}. and saved defaults plan = "free", city = "Pune":

// Campaign row, or a POST /api/v1/call body
{ "contextVariables": { "plan": "pro", "city": "" } }

The agent says: "Calling about your pro plan in Pune." — plan was overridden, and city fell back to the saved default because the row was blank.

Test Inputs values are sent with browser test calls before you save, so you can try a value without committing it. Saving is what turns it into a default for real calls.

End Call

A voice call can end several ways: the agent decides the conversation is done, the caller hangs up, the duration cap is hit, or the silence timeout fires. This section — at the bottom of the Agent tab — configures the endCall tool that owns the graceful path and the farewell spoken on every path.

Every agent has an endCall tool. It is mandatory: if a config is saved without one, the platform auto-seeds a default tool named end_call with the default trigger prompt below. There is no on/off toggle — the LLM can always end the call gracefully.

Fields

FieldTypeDefaultDescription
When to end the callstring (≤ 4,000)See default belowThe trigger prompt — the conditions that should make the LLM call endCall. Becomes the tool's description field. Only contains conditions, not the goodbye text.
Farewell behavior"none", "custom", "audio""custom"What plays before the line goes dead — on every end path: agent hangup, idle timeout, and max duration. The seeded default is custom speaking one of customMessages, which defaults to the single line "Goodbye!"; none also speaks the built-in "Goodbye!".
Farewell messagestring (≤ 4,000)—Required when farewell behavior is custom. The exact text the runner speaks before disconnecting.
Audio recording IDstring (UUID)—Required when farewell behavior is audio. UUID of a pre-recorded file from your org's recordings.

Default trigger prompt

Every agent starts with this prompt. Replace it with your own conditions if your use case has different end-of-conversation cues:

When the conversation is complete, use the end_call tool. A conversation is complete if:
1. The user explicitly says they want to stop (e.g., "That's all," "I'm done," "Goodbye").
2. The user seems satisfied and their goal appears achieved.
3. The user's goal appears achieved based on conversation history, even without explicit confirmation.
Only end the call when you are confident the conversation is truly complete.

How farewell behavior works

The farewell resolves the same way on every end path — the agent calling endCall, the idle-retry limit, or the max-duration cap:

FarewellWhat plays before disconnectWho says goodbye
noneRunner speaks the built-in default, "Goodbye!", waits for playback, then disconnects.The runner — default text.
customRunner picks one of customMessages at random (up to 5 farewell variants), injects it as a TTSSpeakFrame, waits for TTS audio to finish playing, then disconnects. The LLM's own response (if any) is suppressed.The runner — one of your farewell lines.
audioRunner queues the pre-recorded audio file, waits for playback to finish, then disconnects.The pre-recorded voice in the file.

LLM is told to skip the goodbye

The runner adds a system-prompt hint that tells the LLM to call endCall without saying goodbye in its response — because the runner is going to speak the farewell itself. This avoids the double-goodbye where the LLM says "Thanks, goodbye!" and then the runner says "Thanks for calling!" right after.

What ends up in the call config snapshot

Whatever you configure here becomes an entry in the canonical snapshot's tools[] array. After resolution, the snapshot contains:

{
  "tools": [
    {
      "type": "endCall",
      "name": "end_call",
      "description": "<your When-to-end prompt>",
      "config": {
        "messageType":    "custom",
        "customMessages": ["<farewell variant 1>", "<farewell variant 2>"]
      }
    }
  ]
}

If you save or POST a config without any endCall entry, this exact tool is auto-seeded with the default trigger prompt and messageType: "custom" speaking the single farewell "Goodbye!" — a round-trip always returns at least one.

Multiple endCall tools are allowed via the API — give each a distinct name (e.g. end_call and end_call_escalate) and its own trigger description, and the LLM sees them all as separate functions. The first endCall entry is canonical: its delivery config drives the farewell for the automatic end paths (idle timeout, max duration), and this section edits it. Additional entries appear read-only in the Tools tab. See Inline tool format → endCall.

End reason is captured automatically

You do not configure or request an end reason. The runtime sets call.end_reason to exactly one of these enum values based on which exit branch fired:

end_reasonWhen it's setWhat the caller experiences
agent_hangupThe agent's endCall tool fired (configured in this section).Hears the farewell (if any), then line drops.
user_hangupThe caller hung up: WebSocket closed from the client, phone hung up, browser tab closed.They initiated it — they know.
max_durationThe maxCallDurationS cap was hit.Hears the farewell mid-conversation, then the line drops. Set a graceful warning before the cap if this is common.
silence_timeoutThe User Idle handler exhausted userIdleMaxRetries without a response.Heard the idle prompt N times, then the farewell, then the line dropped.
voicemailVoicemail detection classified the first turn as an answering machine (enabled by default in the Call tab).The greeting plays; the agent ends the call instead of talking to the machine.
disconnectUnexpected transport drop, network failure, or pipeline cancellation.Line goes dead unexpectedly.
errorAn unhandled exception inside the pipeline.Line drops; the Call detail page Logs tab shows the traceback.

These enum values map 1:1 onto the public endedReason field on the call row and webhooks (agent_hangup, user_hangup, max_duration, silence_timeout, voicemail; disconnect surfaces as cancelled, error as pipeline_error) — see endedReason values.

Tune the trigger prompt for production

Graceful ending is always available — what's worth tuning is the trigger prompt. If your call history shows a lot of silence_timeout and max_duration reasons, the LLM isn't recognizing when conversations are done: sharpen the end-of-conversation cues in the prompt. Long tails of dead air cost you LLM/TTS/STT the whole time.

Programmatic equivalent

Setting the trigger prompt, farewell custom, and message Thanks for calling — goodbye! in this section produces the same snapshot as POSTing this endCall block via the saved-tool API:

curl -X POST https://dashboard.zoxa.ai/api/v1/tools \
  -H "X-API-Key: zsk_..." \
  -H "Content-Type: application/json" \
  -d '{
    "name": "End call when done",
    "description": "<your trigger conditions>",
    "category": "endCall",
    "definition": {
      "schema_version": 1,
      "type": "endCall",
      "config": {
        "messageType":   "custom",
        "customMessage": "Thanks for calling — goodbye!"
      }
    }
  }'

Or directly inline on a transient POST /api/v1/call — see inline endCall.


Tools Tab

Attach Tools

The Tools tab attaches saved tools to extend your agent's capabilities during calls — looking up data, booking appointments, calling external APIs. Use the Add Tools button in the section header.

FieldTypeDefaultDescription
Toolsstring[] (UUIDs)[]Select from tools you have created in the Tools section. Max 100 tools per agent. Config path: config.tools[].

The config.tools[] array holds a mix of UUID string references (saved tools you select here) and inline tool objects (the mandatory endCall entry configured on the Agent tab, plus any function/query/transferCall tools added via the API). Inline tools created through the API — including any additional endCall tools beyond the first — are shown read-only here so they survive edits. See Tools → tool types and inline tool format.


Engine Tab

The Engine tab holds every model-specific setting plus the low-level pipeline behavior. The middle tabs never select a model — the right-panel chips do — so the Engine tab's Transcriber, Voice, and Model sections always follow whatever the chips have selected. Their available knobs change with the provider.

Sections, top to bottom:

Turn Taking

How the agent decides you've finished speaking, and when you may barge in.

Turn detection is always provider-native: each transcriber's own turn model decides when you've finished speaking, and there is nothing to configure. Tune end-of-turn behavior with each provider's own dials in the transcriber section (uniform 2 s silence cap by default).

The one turn-taking setting is barge-in:

FieldTypeDefaultDescription
Min Words (interruption)integer0Words required to confirm a barge-in while the agent is speaking. 0 = any detected speech interrupts instantly (the provider's own barge-in signal). Config path: stt.interruptionMinWords.

How Min Words ≥ 1 works (provisional interruption). The instant voice activity is detected while the agent is speaking, the agent's audio pauses — so it never talks over the caller. If the transcriber then returns at least Min Words words within about 1.5 s, the interruption is confirmed and the agent stops for real. If it was just a cough, background noise, or a short backchannel ("yeah", "ok"), the agent's audio simply resumes where it left off.

Min WordsFilters noise / "yeah, ok" while agent speaksSingle-word "stop" interrupts the agentOne-word answers when agent is silent
0 (default)❌✅✅
1noise only✅✅
2✅❌ (needs 2 words)✅
3+✅❌✅

Noise Suppression

Voice isolation applied to the caller's audio before transcription — a single on/off card in the editor, on by default. When on, the engine isolates the nearest speaker and removes background voices at full strength.

Turning the toggle on reveals two knobs, mirrored in the API (noiseGate.model — voice-focus or noise-suppression — and noiseGate.enhancementLevel, 0–1, default 0.8). Voice focus isolates the nearest speaker and removes background voices; noise suppression is general background cleanup. See the Agent Config schema.

Transcriber Settings

Advanced tuning for the currently selected STT provider + model (the Transcriber chip). The available knobs differ per provider — for example:

ProviderRepresentative knobs
SonioxcontextTerms, endpoint sensitivity / latency-adjustment level / max-endpoint-delay
Deepgram FluxeotThreshold, eotTimeoutMs, keyterms
AssemblyAImode (balanced / min-latency / max-accuracy), keyterms, voice focus (off / near-field / far-field + strength)
Cartesiaturn start / end thresholds, turnEndTimeoutMs, keyterms
Sarvammode (transcribe / verbatim / translit / codemix), high sensitivity

Full per-provider tables live on the STT Providers page.

Voice Settings

Advanced tuning for the currently selected TTS provider + model (the Voice chip). Examples:

ProviderRepresentative knobs
ElevenLabsstability, similarity, style, speaker boost, speed, text normalization
Cartesiaspeed, volume
Smallestspeed
Sarvampace, temperature (v3)
xAIvoice, speed, streaming latency, text normalization
Inworldspeaking rate, temperature

Full per-provider tables live on the TTS Providers page.

Model Settings

Exactly two dials for the currently selected LLM provider + model (the Model chip): temperature (default 1.0) and max output tokens (default 251). Both are always-on sliders — there is nothing to enable or disable, and no other sampling knobs exist.

Prewarm — the throwaway request that warms the model the instant a call starts so the first reply is as quick as the rest — is always on behind the scenes; there is no switch for it.

Anthropic temperature cap: Anthropic's Messages API accepts temperatures only in the 0.0 -- 1.0 range. When an Anthropic model is selected, the slider caps at 1.0 and any stored value above 1.0 is clamped on save. OpenAI, Google, and other providers accept the full 0.0 -- 2.0 range. A few OpenAI models lock temperature entirely (the gpt-5 family and reasoning models) — the slider is hidden for those.

See LLM Providers for the model catalog and per-model metadata.


Call Tab

The Call tab configures call duration limits, idle behavior, and the human-feel audio layer (background sound, task sound, acknowledge sounds, filler words).

Call Limits

FieldTypeDefaultRangeDescription
Max Call Durationinteger (seconds)61010 -- 7,200, step 10Maximum call length in seconds. The call ends automatically when reached, regardless of conversation state. Displayed as mm:ss in the UI.
User Idle Timeoutinteger (seconds)73 -- 25, step 1How long to wait with no user speech after the agent finishes before triggering an idle prompt.
Idle Promptsstring[] (userIdleMessages)["Are you still there?"]max 5 entriesThe re-engage lines, spoken in order — retry 1 speaks the first entry, retry 2 the second, and so on, letting you escalate naturally ("Are you still there?" → "I can't hear you — are you with me?" → "I'll have to end the call"). If there are more retries than entries, the extra retries speak the built-in default "Are you still there?".
Max Idle Retriesinteger21 -- 5, step 1How many idle prompts play before ending the call. After all retries with no response, the call disconnects with endedReason = "silence_timeout" (on the call row and in webhooks).

Where the farewell lives now

There is no farewell field on the Call tab. The goodbye the agent speaks before hanging up is configured once on the End Call tool (Agent tab) and plays on every end path — agent hangup, idle timeout, and max duration. Leave the End Call farewell behavior on none for the built-in "Goodbye!".

Tip: For production agents, set max call duration based on your expected call length plus a buffer. A 10-minute appointment booking call might use 900 seconds (15 minutes) as the limit. Keep idle retries at 2 -- 3 to avoid annoying users who may have stepped away briefly.

Voicemail Detection

FieldTypeDefaultDescription
Voicemail DetectionbooleanonDetects answering machines on the first turn of an outbound call. When a voicemail greeting is detected, the call ends with endedReason = "voicemail" instead of talking to a machine. For campaign calls, it also triggers a retry when the campaign's retry_on_voicemail toggle is on. Runs a lightweight classifier on the call's own LLM for the first turn only. On by default; turn off for inbound-only agents where a human always answers.

Recording

FieldTypeDefaultDescription
Carrier RecordingbooleanoffTelephony calls only. Also records the call on the telephony provider's side (Twilio/Vobiz console) via the provider's recording API, giving you a second copy in the carrier's Recordings tab. Zoxa's own call recording always runs regardless, so the carrier-side copy is opt-in (some carriers bill per recording).

Background Sound

A faint, continuous environment sound mixed under the whole call — it removes the tell-tale digital silence of a bot line. Off by default; knobs appear when enabled. Config key: backgroundSound.

FieldTypeDefaultRangeDescription
Enabledbooleanoff—Turn the ambient bed on.
Environment (sound)string"library"crowd | library | restaurant | supermarketWhich ambience plays.
Volumenumber0.080.02 -- 0.20Deliberately capped low — above the top of the range ambience reads as noise, not atmosphere.

Task Sound

Keyboard typing or page-turning while the agent runs tools and lookups — like a human agent working at their desk. Purely an audio layer: what the agent says before a tool runs is owned by the tool's own messages. If the tool speaks one, the sound starts after that line finishes (plus the start delay); for silent tools it starts straight into the dead air. Either way it stops the instant the agent starts speaking again. Off by default. Config key: taskSound.

FieldTypeDefaultRangeDescription
Enabledbooleanoff—Turn task sounds on.
Soundstring"typing"typing | page-turnWhich work sound loops during tool execution.
Volumenumber0.350.05 -- 0.60Loudness of the work sound.
Start Delay (startAfterMs)integer (ms)8000 -- 5,000Wait before the sound starts — after the tool's spoken message finishes, or straight into the silence for tools without one. Fast tools finish before it ever plays.

Acknowledge Sounds

Short acknowledgements ("hmm", "okay") in the agent's own voice while the caller speaks at length — signals active listening the way a human agent does. The words are captured once per call in the agent's real TTS voice and mixed directly into the audio, so they never appear in transcripts and can never interrupt or be interrupted by real speech. Off by default. Config key: acknowledgeSounds.

FieldTypeDefaultRangeDescription
Enabledbooleanoff—Turn acknowledge sounds on.
Wordsstring"hmm, hm hmm, okay"comma-separatedThe acknowledgement vocabulary.
Word Order (wordOrder)string"round-robin"round-robin | randomround-robin cycles the words in order (skipped chances don't advance the sequence); random may repeat.
Speech Threshold (thresholdS)number (seconds)32 -- 15How long the caller must keep talking before each acknowledgement chance. The clock measures a continuous talking stretch — it does not reset on brief pauses, only when the agent actually gets a response through.
Probabilitynumber0.80.05 -- 1.0Chance each threshold fires. The skipped chances are what keep the cadence from sounding rhythmic.
Volumenumber1.00.1 -- 1.5Clip loudness relative to the agent's voice.

Filler Words

Natural speech fillers over response latency — the trick human agents (and leading voice assistants) use to cover the thinking gap. If the answer hasn't started within the threshold after the caller finishes speaking, the agent instantly says a filler word ("so", "yeah") in its own voice. The word does appear in the transcript — as the spoken prefix of the answer it precedes, exactly as heard. Off by default. Config key: fillerWords.

FieldTypeDefaultRangeDescription
Enabledbooleanoff—Turn filler words on.
Wordsstring"so, yeah, okay"comma-separatedThe filler vocabulary, captured once per call in the agent's real voice.
Word Order (wordOrder)string"round-robin"round-robin | randomSame semantics as acknowledge sounds.
Silence Threshold (thresholdMs)integer (ms)300100 -- 3,000How long the response may lag before a filler fires. Fires instantly at the threshold — zero synthesis latency (the clip is pre-captured).
Probabilitynumber0.750.05 -- 1.0Chance a lagging response gets a filler.

There is deliberately no volume knob: the filler is the spoken prefix of the answer the voice then continues, so it always plays at exactly the agent's voice level — any offset would make the seam between the filler and the answer audible.

Transcripts and recordings

Acknowledge-sound words never appear in transcripts; filler words do (as the answer's prefix). All four features mix audio at the transport output, after Zoxa's call recording is captured — so recordings contain the agent's clean voice without the ambience/effects.


Analytics Tab

The Analytics tab configures post-call processing: webhooks and conversation summarization.

Post Call Tasks

FieldTypeDefaultDescription
Webhook URLstring or nullnullURL to receive a POST with all call execution data (transcript, duration, extracted variables, etc.) after the call ends. Max 2,048 characters.
Headers (JSON)object or nullnullCustom headers to include in webhook requests. Entered as JSON. Only visible when a webhook URL is set.
Webhook secretstring (write-only)--Used to HMAC-sign webhook deliveries so your endpoint can verify authenticity. Write-only — you can set or replace it, but it is never displayed back (the field shows a "set / replace" affordance, not the current value). Sent as webhookSecret on the agent create/update body; never returned in any API response.
SummarizationbooleanfalseWhen enabled, automatically generates a summary of the conversation after the call ends.

Testing Your Agent

The right panel is a live test surface. You can:

  • Browser call -- Click the "Talk to Agent" button or press Alt+T / Option+T to start a test call directly in your browser. The connection uses SmallWebRTC (peer-to-peer WebRTC with a data channel for live transcript events) — no third-party room service. The talk panel streams the conversation as it happens: your interim speech updates in place until it finalizes, and the agent's reply streams in token by token. During the call, mute with Cmd+M / Ctrl+M and end with Cmd+E / Ctrl+E.
  • Telephony mode -- A test-only toggle above the talk panel. When on, the browser test call is fed through the same 8 kHz μ-law audio path a real phone call uses, so you hear exactly what a caller would hear. It is never saved to the agent and is forced off on real telephony calls. See the callout below.
  • Phone call -- Click "Get Call" to open the phone dialog and place a test call to a real phone number via your configured telephony provider (v2 outbound call).
  • Code preview -- Click the code icon to view the full agent configuration as JSON. This is literally the API request body — copy it straight into POST /api/v1/agents.

Telephony mode — what it does

Simulates phone-call audio (8 kHz, μ-law) during browser test calls so you can hear what callers will hear. Test-only — never saved to the agent. It defaults to off, is passed out-of-band as a telephonySim flag on the test-call request (stripped before config validation), and is forced off on any real telephony call.

Save your agent with Cmd+S (Mac) or Ctrl+S (Windows). The editor shows an "Unsaved" badge when you have pending changes.

Keyboard shortcuts summary:

ShortcutAction
Cmd/Ctrl + SSave agent
Alt/Option + TStart browser test call
Cmd/Ctrl + MToggle mute during call
Cmd/Ctrl + EEnd active call

On this page