Knowledge Base
Upload documents so your voice agents can retrieve and reference information during live calls using RAG (Retrieval-Augmented Generation).
What you'll learn
- How to upload documents and choose a retrieval mode
- The difference between Full Document and Chunked Search retrieval
- How RAG search works during a live call
- How to manage documents via the dashboard and API
The Knowledge Base lets you upload documents that your agents can search and reference during calls. When a caller asks a question that requires specific information -- a return policy, product specifications, pricing details -- the agent queries the knowledge base, retrieves the relevant content, and uses it to answer accurately.
This is powered by RAG (Retrieval-Augmented Generation). Documents are processed, embedded, and stored so they can be searched semantically during live conversations.
Supported File Types
| Format | Extension | MIME Type |
|---|---|---|
.pdf | application/pdf | |
| Word Document | .docx | application/vnd.openxmlformats-officedocument.wordprocessingml.document |
| Word Document (legacy) | .doc | application/msword |
| Plain Text | .txt | text/plain |
| JSON | .json | application/json |
| Markdown | .md | text/markdown |
| CSV | .csv | text/csv |
Maximum file size: 5 MB per document.
The 5 MB limit is enforced during background processing, not during upload. Files larger than 5 MB will upload successfully but will fail processing with an error message indicating the size limit was exceeded.
Retrieval Modes
When uploading a document, you choose how the agent should retrieve information from it. This choice is permanent and cannot be changed after upload.
Full Document
The entire document text is extracted and stored. On every retrieval, the complete content is provided to the LLM. No chunking, embedding, or vector search is performed.
| Aspect | Detail |
|---|---|
| Best for | Small reference documents where the agent needs complete context -- menus, price lists, FAQs, short policy documents, product catalogs |
| How it works | Full text is extracted and stored. On retrieval, the entire text is returned with similarity score 1.0 |
| Embedding required | No -- no embeddings API key needed for Full Document mode |
| Trade-off | Uses more LLM context window per retrieval. Not suitable for documents over ~10 pages |
Chunked Search
The document is split into smaller chunks, and each chunk is embedded using the configured embedding model (default: OpenAI text-embedding-3-small, 1536-dimensional vectors). During calls, the agent's query is embedded and compared against chunk embeddings using vector similarity search. Only the most relevant chunks are returned.
| Aspect | Detail |
|---|---|
| Best for | Large documents where only a small section is relevant at a time -- employee handbooks, technical manuals, legal policies, knowledge articles |
| How it works | Text is extracted, split into chunks (128 tokens per chunk), and each chunk is embedded. Retrieval uses vector similarity to find top matches |
| Embedding required | Yes -- handled automatically by the platform using OpenAI text-embedding-3-small |
| Trade-off | Retrieval quality depends on how well the query matches the relevant chunks. Requires more processing time during upload |
Which mode should I use? If the document is short (under 5 pages) and the agent needs all the information in context, use Full Document. If the document is long and only small sections are relevant per query, use Chunked Search.
How Documents Are Processed
When you upload a document, the following happens:
Upload -- the file is uploaded to secure object storage (S3/MinIO) via a presigned URL.
Record creation -- a document record is created in the database with status pending.
Background processing -- an asynchronous background task processes the document:
- For Full Document mode: the full text is extracted (using Docling for PDF/DOCX, direct read for TXT/JSON) and stored on the document record.
- For Chunked mode: the text is extracted, split into chunks using a hybrid chunker, and each chunk is embedded using the configured embedding model.
Status update -- the document status moves to processing, then completed (or failed if an error occurs).
Processing Statuses
| Status | Description |
|---|---|
pending | Document uploaded, waiting to be processed. |
processing | Background task is extracting, chunking, and embedding. |
completed | Document is ready for retrieval during calls. |
failed | Processing encountered an error. Check the error message for details (common causes: missing embeddings API key, file too large, duplicate file, corrupt document). |
The dashboard polls automatically while documents are processing, so you will see the status update in real time.
Uploading Documents
Navigate to Knowledge Base in the dashboard sidebar.
Click Upload Document.
Drag and drop a file or click to browse. Supported formats: .pdf, .docx, .doc, .txt, .json (max 5 MB). (.md and .csv are also accepted via the Files API.)
Select the retrieval mode:
- Full Document -- entire text provided on each retrieval.
- Chunked Search -- split into chunks with vector similarity search.
Click Upload & Process.
The document will appear in your list with a Pending badge, then Processing, then Completed.
Attaching Documents to Agents
After a document is processed, you attach it to an agent to make it available during calls: in the agent editor, go to the Tools tab and select documents under Knowledge Base Documents.
Only documents with status completed appear in the document selector. You can attach multiple documents to a single agent.
Attaching a document does not inject it into the system prompt. Instead, it registers a retrieve_from_knowledge_base function that the LLM can call when it needs to look up information. This means the LLM decides when to search -- it does not receive the entire document on every turn.
How RAG Search Works During Calls
When a tool-registered agent encounters a question that might require document knowledge, the following happens:
LLM triggers retrieval -- the LLM calls the retrieve_from_knowledge_base function with a search query it formulates from the conversation.
Retrieval executes based on the document's mode:
- Full Document documents: the complete text is returned directly (similarity score
1.0). - Chunked documents: the query is embedded, and vector similarity search returns the top matching chunks (default: top 3).
Results returned to LLM -- the retrieved text chunks (with source filenames and similarity scores) are passed back to the LLM.
LLM responds -- the LLM incorporates the retrieved information into its spoken response.
If both Full Document and Chunked documents are attached, full documents are returned in their entirety while chunked documents go through vector search -- both sets of results are combined.
The Retrieval Function Schema
This is what the LLM sees when knowledge base documents are attached:
{
"type": "function",
"function": {
"name": "retrieve_from_knowledge_base",
"description": "Retrieve relevant information from specific documents in the knowledge base. Use this tool when you need to look up facts, policies, procedures, or any information that might be stored in the available documents.",
"parameters": {
"type": "object",
"properties": {
"query": {
"type": "string",
"description": "The search query to find relevant information. Be specific and use natural language. Example: 'What is the refund policy for canceled orders?'"
}
},
"required": ["query"]
}
}
}Embedding Model
All embeddings use OpenAI text-embedding-3-small (1536-dimensional vectors). This is managed by the platform -- no API key configuration is needed on your part.
| Setting | Value | Description |
|---|---|---|
| Provider | OpenAI | Managed by the platform. |
| Model | text-embedding-3-small | 1536-dimensional embeddings. |
| Chunking | MiniLM tokenizer, 128 tokens per chunk | Hybrid chunking with Docling for structured formats (PDF, DOCX) and token-based splitting for plain text. |
The same embedding model is used for both indexing (at upload time) and querying (during calls). This ensures vector compatibility -- embedding models cannot be mixed.
Managing Documents
Listing Documents
The Knowledge Base page shows all documents in your organization with:
- Filename and file size
- Processing status badge
- Retrieval mode badge (Full Document or Chunked)
- Number of chunks (for chunked documents)
- Upload date
- Search/filter by filename
Deleting Documents
Click the delete button on any document to soft-delete it. Deleted documents are no longer available for retrieval during calls. This action removes the document and all its chunks from search results.
Duplicate Detection
If you upload a file that is identical to an existing document (based on file hash), the system detects the duplicate and marks the upload as failed with a message indicating which existing document it duplicates. You must delete the duplicate files and consolidate them before re-uploading.
API Reference
Get Upload URL
Request a presigned URL for uploading a document.
curl -X POST https://dashboard.zoxa.ai/api/v1/knowledge-base/upload-url \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"filename": "company-faq.pdf",
"mime_type": "application/pdf"
}'Response:
{
"upload_url": "https://storage.example.com/presigned-put-url...",
"document_uuid": "a1b2c3d4-e5f6-7890-abcd-ef1234567890",
"s3_key": "knowledge_base/1/a1b2c3d4-e5f6-7890-abcd-ef1234567890/company-faq.pdf"
}Upload the File
Upload the file to the presigned URL using a PUT request:
curl -X PUT "PRESIGNED_UPLOAD_URL" \
-H "Content-Type: application/pdf" \
--data-binary @company-faq.pdfTrigger Processing
After uploading, trigger document processing:
curl -X POST https://dashboard.zoxa.ai/api/v1/knowledge-base/process-document \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"document_uuid": "a1b2c3d4-e5f6-7890-abcd-ef1234567890",
"s3_key": "knowledge_base/1/a1b2c3d4-e5f6-7890-abcd-ef1234567890/company-faq.pdf",
"retrieval_mode": "chunked"
}'The retrieval_mode field accepts:
"chunked"-- vector similarity search (default)"full_document"-- full text retrieval
List Documents
curl "https://dashboard.zoxa.ai/api/v1/knowledge-base/documents?limit=50&offset=0" \
-H "Authorization: Bearer YOUR_API_KEY"Response:
{
"documents": [
{
"id": 1,
"document_uuid": "a1b2c3d4-...",
"filename": "company-faq.pdf",
"file_size_bytes": 245760,
"processing_status": "completed",
"retrieval_mode": "chunked",
"total_chunks": 42,
"created_at": "2026-05-22T10:30:00Z",
"is_active": true
}
],
"total": 1,
"limit": 50,
"offset": 0
}Get Document Details
curl "https://dashboard.zoxa.ai/api/v1/knowledge-base/documents/DOCUMENT_UUID" \
-H "Authorization: Bearer YOUR_API_KEY"Delete Document
curl -X DELETE "https://dashboard.zoxa.ai/api/v1/knowledge-base/documents/DOCUMENT_UUID" \
-H "Authorization: Bearer YOUR_API_KEY"Search Chunks
Search for document chunks similar to a query (useful for testing retrieval quality):
curl -X POST https://dashboard.zoxa.ai/api/v1/knowledge-base/search \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"query": "What is the refund policy?",
"limit": 5,
"document_uuids": ["a1b2c3d4-..."],
"min_similarity": 0.7
}'Response:
{
"chunks": [
{
"id": 123,
"document_id": 1,
"chunk_text": "Our refund policy allows returns within 30 days of purchase...",
"contextualized_text": "Refund Policy: Our refund policy allows returns within 30 days...",
"chunk_index": 7,
"chunk_metadata": {},
"filename": "company-faq.pdf",
"document_uuid": "a1b2c3d4-...",
"similarity": 0.89
}
],
"query": "What is the refund policy?",
"total_results": 1
}| Parameter | Type | Default | Description |
|---|---|---|---|
query | string | Required | Search query text. |
limit | integer | 5 | Maximum results (1-50). |
document_uuids | string[] | null | Filter by specific documents. |
min_similarity | float | null | Minimum similarity threshold (0.0-1.0). |
Files API (Simplified Upload)
For API developers who want a simpler upload experience, the Files API (/api/v1/files) lets you upload a document in a single call -- no presigned URLs, no separate processing trigger.
curl -X POST https://dashboard.zoxa.ai/api/v1/files \
-H "X-API-Key: zsk_..." \
-F [email protected]The response includes an id that you can use immediately as a documentId in inline query tools during transient agent calls.
The Files API and the Knowledge Base dashboard routes (/knowledge-base/*) both create the same underlying documents. The Files API is a convenience wrapper for programmatic access. See the full Files API reference.
Best Practices
- Choose the right retrieval mode -- use Full Document for short reference docs (menus, price lists). Use Chunked Search for anything over 5 pages.
- Keep documents focused -- one topic per document produces better retrieval results than one large document covering everything.
- Test retrieval quality -- use the
/searchAPI endpoint to test queries against your documents before deploying to production agents. - Update documents by re-uploading -- to update a document, delete the old version and upload the new one. Documents are immutable after processing.
- Use descriptive filenames -- filenames are included in retrieval results, helping the LLM identify which source it is citing.
- Use the Files API for automation -- if you are uploading documents programmatically, use
POST /api/v1/filesinstead of the 3-step presigned URL flow.
Knowledge Base Tool
Configure tools that search your uploaded documents during voice calls, giving your agent instant access to product docs, policies, FAQs, and any other reference material via RAG retrieval.
LLM Providers
Every language model provider supported by zoxaAI -- available models, default configuration, and provider-specific options.