Key Concepts
The core building blocks of zoxaAI -- agents, calls, runs, campaigns, tools, knowledge bases, and providers -- explained with examples, tables, and practical context.
In this guide
This page explains every major concept in zoxaAI. Read it end to end to build a solid mental model of the platform, or jump to a specific section when you need a reference.
Understanding these concepts will make everything else in the platform click into place. Each section below covers what the concept is, when you use it, and how it connects to the rest of the system.
Agents
An agent is a voice AI persona. It defines how the AI presents itself during a call -- what it says, how it reasons, what voice it speaks in, and what actions it can take. Every call in zoxaAI is powered by an agent.
Core settings
Every agent has three required settings:
| Setting | What it controls | Example |
|---|---|---|
| System prompt | The instructions that tell the language model how to behave -- role, personality, knowledge, boundaries, response style. | "You are a support agent for Acme Inc. Help callers reset their passwords..." |
| Voice | The TTS voice the caller hears. Each voice belongs to a TTS provider. | ElevenLabs Rachel, Cartesia Sophia, Inworld Ashley |
| Language model | The AI model that generates responses during the conversation. | GPT-5.4-mini, Claude Sonnet 4.6, Gemini 2.5 Flash |
Beyond these three, agents have extensive configuration across five editor tabs (Agent, Engine, Tools, Call, Analytics) covering welcome messages, interruption behavior, silence detection, speech speed, temperature, tool attachments, call duration limits, and more. See the full Agent configuration reference.
Calls
A call is the real-time audio connection between a person and your agent. zoxaAI supports three call types, each suited to a different scenario.
| Type | Transport | How it starts | Latency | Phone number required | Typical use case |
|---|---|---|---|---|---|
| Web Call | WebRTC (browser) | User clicks "Web Call" in the dashboard | Lowest (no telephony hop) | No | Testing, in-app voice support |
| Phone Call | Telephony (Twilio, Vonage, etc.) | Inbound: caller dials your connected phone number. Outbound: you trigger via the dashboard or API. | Low (telephony adds ~200ms) | Yes | Production support lines, outbound sales, appointment reminders |
| Campaign Call | Telephony | zoxaAI dials automatically as part of a batch campaign | Low | Yes | Outbound campaigns: surveys, reminders, sales cadences |
All three call types produce the same run record when they end, so your analytics, transcripts, and recordings are consistent regardless of how the call was initiated.
How to start a call
| Method | Call type | Details |
|---|---|---|
| Dashboard Web Call button | Web Call | Click "Web Call" on any agent page. Uses your browser microphone. |
| Inbound phone number | Phone Call | A caller dials your connected number. The agent answers automatically. |
| Dashboard outbound call | Phone Call | Enter a phone number in the dashboard and click "Call". |
API POST /call | Phone Call | Trigger an outbound call programmatically with type: "outbound". |
| Campaign | Campaign Call | Upload a CSV of contacts and start the campaign. zoxaAI dials each contact. |
Web calls are free of telephony charges
Web calls use WebRTC directly from the browser, so there are no telephony provider charges -- you only pay for STT, LLM, and TTS usage. This makes web calls ideal for testing and for use cases where your users are already on your website or app.
Runs
A run is the complete record of a single call. Every call -- regardless of type -- creates a run when it ends. Runs are the primary unit of data in zoxaAI: they contain everything you need to evaluate agent performance, debug issues, and extract business value.
Run record fields
| Field | Type | Description |
|---|---|---|
| Call ID | string | Unique identifier for the run. |
| Agent | string | Which agent handled the call. |
| Transport | string | How the call connected: webrtc, outbound, or inbound. |
| Status | string | How the call ended (see status table below). |
| Transcript | array | Full text of the conversation, speaker-labeled (AI and Human), with timestamps per turn. |
| Recording | url | Audio file of the complete call. Downloadable from the dashboard or API. |
| Duration | float | Length of the call in seconds. |
| Cost | float | Total cost of the run in USD. |
| Cost breakdown | object | Itemized cost by component: LLM tokens, TTS characters, STT minutes, telephony minutes. |
| Platform number | string | Your phone number (caller ID for outbound, inbound number for inbound). Null for web calls. |
| Customer number | string | The other party's phone number. Null for web calls. |
| Telephony provider | string | Which provider handled the call (twilio, vonage, etc.). Null for web calls. |
| Tags | array | Labels like outbound, inbound, campaign, agent. |
| Campaign | string | The campaign that initiated the call, if applicable. |
| Ended reason | string | Specific reason the call ended (e.g., completed, hung_up, error, max_duration). |
| Started at | datetime | When the call started. |
| Ended at | datetime | When the call ended. |
Run statuses
| Status | Meaning | What happened |
|---|---|---|
started | In progress | The call is currently active. This status is temporary. |
completed | Finished normally | The conversation reached its natural end. The agent or caller ended the call gracefully. |
failed | Error | An error occurred that ended the call unexpectedly -- pipeline failure, provider timeout, etc. |
busy | Busy signal | The number dialed returned a busy signal. Only applies to outbound phone/campaign calls. |
no_answer | Not picked up | The call was not answered within the timeout window. Only applies to outbound phone/campaign calls. |
Where to find runs
- Agent Runs tab: Shows runs for a specific agent.
- Calls page (sidebar): Shows all runs across all agents. Supports filtering by status, transport, date range, and tags. See Call History.
- API: Query runs programmatically via the API.
Campaigns
A campaign is a batch of outbound calls made to a list of contacts. Instead of triggering calls one by one, you upload a CSV of phone numbers, assign an agent, and zoxaAI works through the list automatically.
Campaign lifecycle
Create the campaign
Go to Campaigns in the sidebar and click Create New Campaign. Set a name, select the target agent, and upload a CSV file.
The CSV must have a phone_number column (E.164 format recommended). Any additional columns become context variables available during the call:
phone_number,first_name,account_id,appointment_date
+14155552671,Sarah,ACC-1234,2026-06-15
+14155552672,James,ACC-5678,2026-06-16In this example, first_name, account_id, and appointment_date are injected into the agent's context so it can personalize the conversation (e.g., "Hi Sarah, I'm calling about your appointment on June 15th").
Configure settings
Optionally set advanced parameters:
| Setting | What it controls | Default |
|---|---|---|
| Telephony configuration | Which provider account + phone-number pool the campaign dials from | Org default |
| Concurrency | How many calls run in parallel | Platform default |
| Retries | How many times to reattempt busy or no_answer calls | Configurable |
| Retry delay | Wait time between retry attempts | Configurable |
| Schedule | Start time and calling hours window | Immediate |
| Circuit breaker | Auto-pause the campaign if failure rate exceeds a threshold | Enabled |
The from-number pool is keyed per telephony configuration, so two campaigns running on different configs never share or steal each other's outbound caller IDs.
Start the campaign
Click Start. The campaign moves through these states:
| State | Description |
|---|---|
created | Configured but not started. You can still edit settings. |
syncing | CSV data is being processed and contacts are being queued. |
running | Actively dialing contacts. |
paused | Temporarily stopped. Can be resumed. Also triggered automatically by circuit breaker. |
completed | All contacts have been processed. |
failed | Critical error occurred. |
Monitor and review
Each contact call produces its own run record. You can monitor progress on the campaign detail page and review individual call transcripts and recordings. The detail page also shows a structured event log for the campaign — circuit-breaker trips (with the last 20 failures attached), phone-pool exhaustion retries, and CSV sync issues all show up there.
Circuit breaker protection
If an unusual number of calls fail in a short period (e.g., wrong numbers, provider outages), the circuit breaker automatically pauses the campaign. This prevents runaway errors and wasted spend. The trip event includes the last 20 failed runs so you can see exactly what went wrong — review the failures and resume manually.
Tools
Tools give your agent the ability to take actions during a call, beyond just speaking. When the LLM determines that a tool is needed based on the conversation context and your instructions, it triggers the tool automatically.
| Tool | What it does | Example |
|---|---|---|
| HTTP API | Makes an HTTP request (GET, POST, PUT, PATCH, DELETE) to any URL during the call. Values extracted from the conversation are injected into the request as parameters. | Check appointment availability, look up account data, submit a form, create a CRM record. |
| Call Transfer | Transfers the caller to a phone number -- for example, a human agent, a department extension, or a payment line. | "Let me transfer you to our billing department." |
| End Call | Ends the call programmatically when a condition is met. | End the call after the agent has collected all required information and confirmed with the caller. |
| Knowledge Base | Searches uploaded documents using RAG and returns relevant content to the LLM. | "What is your refund policy?" — agent searches the uploaded policy document and answers accurately. |
How tools work
- You define tools with a name, description, and parameters (for HTTP API tools).
- You attach them to an agent.
- During the call, the LLM reads the tool name and description to decide when to use it.
- When triggered, the platform executes the tool (e.g., sends the HTTP request) and returns the result to the LLM.
- The LLM incorporates the result into its next response.
For detailed configuration, see the Tools reference.
Knowledge Base
A knowledge base lets your agent look up information during a call using retrieval-augmented generation (RAG). You upload documents, zoxaAI indexes them, and when a caller asks something the agent does not know from its system prompt alone, the agent searches the knowledge base and uses the retrieved content to answer accurately.
Supported file types
| Format | Extension | Max size |
|---|---|---|
.pdf | 5 MB | |
| Word Document | .docx, .doc | 5 MB |
| Plain Text | .txt | 5 MB |
| JSON | .json | 5 MB |
| Markdown | .md | 5 MB |
| CSV | .csv | 5 MB |
Retrieval modes
When you upload a document, you choose one of two retrieval modes. This cannot be changed after upload.
| Mode | How it works | Best for |
|---|---|---|
| Full Document | The entire document text is provided to the LLM on every retrieval. No chunking or embedding search. | Small reference docs: menus, price lists, FAQs, short policy documents, product catalogs. |
| Chunked Search | The document is split into chunks, each embedded using text-embedding-3-small. During calls, the agent's query is embedded and compared against chunks using vector similarity. Only the most relevant chunks are returned. | Large documents: employee handbooks, technical manuals, legal policies, knowledge articles. |
How it works during a call
- The caller asks a question the agent cannot answer from its system prompt alone.
- The LLM generates a search query based on the conversation context.
- zoxaAI searches the knowledge base and retrieves relevant content (full document or matching chunks).
- The retrieved content is injected into the LLM context.
- The LLM uses the content to generate an accurate, grounded response.
Shared across agents
A single knowledge base can be attached to multiple agents. You can add, remove, or replace documents at any time -- updates take effect on the next call.
For full configuration details, see the Knowledge Base reference.
Providers
zoxaAI provides managed access to a wide range of AI and telephony providers. You do not need to create separate accounts or manage API keys -- everything is included with your zoxaAI account.
Models are a curated catalog — you pick from the models each provider offers below. Every model carries latency, cost, and quality metadata, and off-catalog model ids are rejected at save time.
LLM providers (4)
| Provider | Example models |
|---|---|
| OpenAI | GPT-5.4, GPT-5.4-mini, GPT-5.4-nano, GPT-6 Luna, GPT-5.2, GPT-5.1, GPT-4.1, GPT-4.1-mini |
| Google (Gemini) | Gemini 2.5 Flash, Gemini 2.5 Flash-Lite, Gemini 3.5 Flash-Lite, Gemini 3.1 Flash-Lite, Gemini 3 Flash Preview |
| Anthropic (Claude) | Claude Sonnet 4.6, Claude Haiku 4.5 |
| Qwen (Alibaba) | qwen-flash, qwen3.6-flash, qwen3.7-plus |
TTS providers (6)
| Provider | Notable features |
|---|---|
| ElevenLabs | Multilingual, voice cloning, stability/similarity controls |
| Cartesia | Sonic 3.6, speed/volume controls, low latency |
| Smallest | Lightning v3.1, strong Indian language support |
| Sarvam | Bulbul v3, Indian language specialist |
| Inworld | Inworld TTS 2, expressive voices, 22 languages |
| xAI | Grok Voice TTS |
STT providers (5)
| Provider | Notable features |
|---|---|
| Soniox | Real-time v5, 56+ languages, native code-switching across the agent's languages list |
| Deepgram Flux | Flux general models, server-side end-of-turn detection, keyterm boosting |
| AssemblyAI | Universal-3.6 Pro Realtime, high accuracy with native code-switching |
| Cartesia | Ink-2, server-side turn detection |
| Sarvam | Saaras v3, Indian language specialist, code-mix |
Telephony providers (8)
| Provider | Type |
|---|---|
| Twilio | Cloud telephony |
| Vonage | Cloud telephony |
| Telnyx | Cloud telephony |
| Plivo | Cloud telephony |
| Cloudonix | Cloud telephony |
| VoBiz | Cloud telephony |
| Tata Smartflo | Cloud telephony (streaming-native) |
| Asterisk ARI | Self-hosted PBX |
For detailed provider configuration and model lists, see:
LLM Providers
All 4 providers, models, and configuration options
TTS Providers
All 6 providers, voices, speed/quality settings
STT Providers
All 5 providers, models, language support
How everything connects
Here is how the core concepts relate to each other in practice:
Agent
├── System prompt, voice, LLM model
├── Tools (HTTP API, Call Transfer, End Call)
├── Knowledge Base (optional, shared across agents)
│
├── Receives a Call (Web, Phone, or Campaign)
│ └── Runs the STT → LLM → TTS pipeline in real time
│ └── Uses tools and knowledge base as needed
│
└── Produces a Run when the call ends
├── Transcript, recording, duration, cost
└── Accessible via dashboard or APIA Campaign is an orchestration layer on top of calls -- it takes a list of contacts and triggers one call (and therefore one run) per contact, with concurrency, retries, and scheduling handled automatically.