zoxaAI
Homepage

Knowledge Base

Upload documents so your voice agents can retrieve and reference information during live calls using RAG (Retrieval-Augmented Generation).

What you'll learn

  • How to upload documents and choose a retrieval mode
  • The difference between Full Document and Chunked Search retrieval
  • How RAG search works during a live call
  • How to manage documents via the dashboard and API

The Knowledge Base lets you upload documents that your agents can search and reference during calls. When a caller asks a question that requires specific information -- a return policy, product specifications, pricing details -- the agent queries the knowledge base, retrieves the relevant content, and uses it to answer accurately.

This is powered by RAG (Retrieval-Augmented Generation). Documents are processed, embedded, and stored so they can be searched semantically during live conversations.


Supported File Types

FormatExtensionMIME Type
PDF.pdfapplication/pdf
Word Document.docxapplication/vnd.openxmlformats-officedocument.wordprocessingml.document
Word Document (legacy).docapplication/msword
Plain Text.txttext/plain
JSON.jsonapplication/json
Markdown.mdtext/markdown
CSV.csvtext/csv

Maximum file size: 5 MB per document.

The 5 MB limit is enforced during background processing, not during upload. Files larger than 5 MB will upload successfully but will fail processing with an error message indicating the size limit was exceeded.


Retrieval Modes

When uploading a document, you choose how the agent should retrieve information from it. This choice is permanent and cannot be changed after upload.

Full Document

The entire document text is extracted and stored. On every retrieval, the complete content is provided to the LLM. No chunking, embedding, or vector search is performed.

AspectDetail
Best forSmall reference documents where the agent needs complete context -- menus, price lists, FAQs, short policy documents, product catalogs
How it worksFull text is extracted and stored. On retrieval, the entire text is returned with similarity score 1.0
Embedding requiredNo -- no embeddings API key needed for Full Document mode
Trade-offUses more LLM context window per retrieval. Not suitable for documents over ~10 pages

The document is split into smaller chunks, and each chunk is embedded using the configured embedding model (default: OpenAI text-embedding-3-small, 1536-dimensional vectors). During calls, the agent's query is embedded and compared against chunk embeddings using vector similarity search. Only the most relevant chunks are returned.

AspectDetail
Best forLarge documents where only a small section is relevant at a time -- employee handbooks, technical manuals, legal policies, knowledge articles
How it worksText is extracted, split into chunks (128 tokens per chunk), and each chunk is embedded. Retrieval uses vector similarity to find top matches
Embedding requiredYes -- handled automatically by the platform using OpenAI text-embedding-3-small
Trade-offRetrieval quality depends on how well the query matches the relevant chunks. Requires more processing time during upload

Which mode should I use? If the document is short (under 5 pages) and the agent needs all the information in context, use Full Document. If the document is long and only small sections are relevant per query, use Chunked Search.


How Documents Are Processed

When you upload a document, the following happens:

Upload -- the file is uploaded to secure object storage (S3/MinIO) via a presigned URL.

Record creation -- a document record is created in the database with status pending.

Background processing -- an asynchronous background task processes the document:

  • For Full Document mode: the full text is extracted (using Docling for PDF/DOCX, direct read for TXT/JSON) and stored on the document record.
  • For Chunked mode: the text is extracted, split into chunks using a hybrid chunker, and each chunk is embedded using the configured embedding model.

Status update -- the document status moves to processing, then completed (or failed if an error occurs).

Processing Statuses

StatusDescription
pendingDocument uploaded, waiting to be processed.
processingBackground task is extracting, chunking, and embedding.
completedDocument is ready for retrieval during calls.
failedProcessing encountered an error. Check the error message for details (common causes: missing embeddings API key, file too large, duplicate file, corrupt document).

The dashboard polls automatically while documents are processing, so you will see the status update in real time.


Uploading Documents

Navigate to Knowledge Base in the dashboard sidebar.

Click Upload Document.

Drag and drop a file or click to browse. Supported formats: .pdf, .docx, .doc, .txt, .json (max 5 MB). (.md and .csv are also accepted via the Files API.)

Select the retrieval mode:

  • Full Document -- entire text provided on each retrieval.
  • Chunked Search -- split into chunks with vector similarity search.

Click Upload & Process.

The document will appear in your list with a Pending badge, then Processing, then Completed.


Attaching Documents to Agents

After a document is processed, you attach it to an agent to make it available during calls: in the agent editor, go to the Tools tab and select documents under Knowledge Base Documents.

Only documents with status completed appear in the document selector. You can attach multiple documents to a single agent.

Attaching a document does not inject it into the system prompt. Instead, it registers a retrieve_from_knowledge_base function that the LLM can call when it needs to look up information. This means the LLM decides when to search -- it does not receive the entire document on every turn.


How RAG Search Works During Calls

When a tool-registered agent encounters a question that might require document knowledge, the following happens:

LLM triggers retrieval -- the LLM calls the retrieve_from_knowledge_base function with a search query it formulates from the conversation.

Retrieval executes based on the document's mode:

  • Full Document documents: the complete text is returned directly (similarity score 1.0).
  • Chunked documents: the query is embedded, and vector similarity search returns the top matching chunks (default: top 3).

Results returned to LLM -- the retrieved text chunks (with source filenames and similarity scores) are passed back to the LLM.

LLM responds -- the LLM incorporates the retrieved information into its spoken response.

If both Full Document and Chunked documents are attached, full documents are returned in their entirety while chunked documents go through vector search -- both sets of results are combined.

The Retrieval Function Schema

This is what the LLM sees when knowledge base documents are attached:

{
  "type": "function",
  "function": {
    "name": "retrieve_from_knowledge_base",
    "description": "Retrieve relevant information from specific documents in the knowledge base. Use this tool when you need to look up facts, policies, procedures, or any information that might be stored in the available documents.",
    "parameters": {
      "type": "object",
      "properties": {
        "query": {
          "type": "string",
          "description": "The search query to find relevant information. Be specific and use natural language. Example: 'What is the refund policy for canceled orders?'"
        }
      },
      "required": ["query"]
    }
  }
}

Embedding Model

All embeddings use OpenAI text-embedding-3-small (1536-dimensional vectors). This is managed by the platform -- no API key configuration is needed on your part.

SettingValueDescription
ProviderOpenAIManaged by the platform.
Modeltext-embedding-3-small1536-dimensional embeddings.
ChunkingMiniLM tokenizer, 128 tokens per chunkHybrid chunking with Docling for structured formats (PDF, DOCX) and token-based splitting for plain text.

The same embedding model is used for both indexing (at upload time) and querying (during calls). This ensures vector compatibility -- embedding models cannot be mixed.


Managing Documents

Listing Documents

The Knowledge Base page shows all documents in your organization with:

  • Filename and file size
  • Processing status badge
  • Retrieval mode badge (Full Document or Chunked)
  • Number of chunks (for chunked documents)
  • Upload date
  • Search/filter by filename

Deleting Documents

Click the delete button on any document to soft-delete it. Deleted documents are no longer available for retrieval during calls. This action removes the document and all its chunks from search results.

Duplicate Detection

If you upload a file that is identical to an existing document (based on file hash), the system detects the duplicate and marks the upload as failed with a message indicating which existing document it duplicates. You must delete the duplicate files and consolidate them before re-uploading.


API Reference

Get Upload URL

Request a presigned URL for uploading a document.

curl -X POST https://dashboard.zoxa.ai/api/v1/knowledge-base/upload-url \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "filename": "company-faq.pdf",
    "mime_type": "application/pdf"
  }'

Response:

{
  "upload_url": "https://storage.example.com/presigned-put-url...",
  "document_uuid": "a1b2c3d4-e5f6-7890-abcd-ef1234567890",
  "s3_key": "knowledge_base/1/a1b2c3d4-e5f6-7890-abcd-ef1234567890/company-faq.pdf"
}

Upload the File

Upload the file to the presigned URL using a PUT request:

curl -X PUT "PRESIGNED_UPLOAD_URL" \
  -H "Content-Type: application/pdf" \
  --data-binary @company-faq.pdf

Trigger Processing

After uploading, trigger document processing:

curl -X POST https://dashboard.zoxa.ai/api/v1/knowledge-base/process-document \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "document_uuid": "a1b2c3d4-e5f6-7890-abcd-ef1234567890",
    "s3_key": "knowledge_base/1/a1b2c3d4-e5f6-7890-abcd-ef1234567890/company-faq.pdf",
    "retrieval_mode": "chunked"
  }'

The retrieval_mode field accepts:

  • "chunked" -- vector similarity search (default)
  • "full_document" -- full text retrieval

List Documents

curl "https://dashboard.zoxa.ai/api/v1/knowledge-base/documents?limit=50&offset=0" \
  -H "Authorization: Bearer YOUR_API_KEY"

Response:

{
  "documents": [
    {
      "id": 1,
      "document_uuid": "a1b2c3d4-...",
      "filename": "company-faq.pdf",
      "file_size_bytes": 245760,
      "processing_status": "completed",
      "retrieval_mode": "chunked",
      "total_chunks": 42,
      "created_at": "2026-05-22T10:30:00Z",
      "is_active": true
    }
  ],
  "total": 1,
  "limit": 50,
  "offset": 0
}

Get Document Details

curl "https://dashboard.zoxa.ai/api/v1/knowledge-base/documents/DOCUMENT_UUID" \
  -H "Authorization: Bearer YOUR_API_KEY"

Delete Document

curl -X DELETE "https://dashboard.zoxa.ai/api/v1/knowledge-base/documents/DOCUMENT_UUID" \
  -H "Authorization: Bearer YOUR_API_KEY"

Search Chunks

Search for document chunks similar to a query (useful for testing retrieval quality):

curl -X POST https://dashboard.zoxa.ai/api/v1/knowledge-base/search \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "query": "What is the refund policy?",
    "limit": 5,
    "document_uuids": ["a1b2c3d4-..."],
    "min_similarity": 0.7
  }'

Response:

{
  "chunks": [
    {
      "id": 123,
      "document_id": 1,
      "chunk_text": "Our refund policy allows returns within 30 days of purchase...",
      "contextualized_text": "Refund Policy: Our refund policy allows returns within 30 days...",
      "chunk_index": 7,
      "chunk_metadata": {},
      "filename": "company-faq.pdf",
      "document_uuid": "a1b2c3d4-...",
      "similarity": 0.89
    }
  ],
  "query": "What is the refund policy?",
  "total_results": 1
}
ParameterTypeDefaultDescription
querystringRequiredSearch query text.
limitinteger5Maximum results (1-50).
document_uuidsstring[]nullFilter by specific documents.
min_similarityfloatnullMinimum similarity threshold (0.0-1.0).

Files API (Simplified Upload)

For API developers who want a simpler upload experience, the Files API (/api/v1/files) lets you upload a document in a single call -- no presigned URLs, no separate processing trigger.

curl -X POST https://dashboard.zoxa.ai/api/v1/files \
  -H "X-API-Key: zsk_..." \
  -F [email protected]

The response includes an id that you can use immediately as a documentId in inline query tools during transient agent calls.

The Files API and the Knowledge Base dashboard routes (/knowledge-base/*) both create the same underlying documents. The Files API is a convenience wrapper for programmatic access. See the full Files API reference.


Best Practices

  1. Choose the right retrieval mode -- use Full Document for short reference docs (menus, price lists). Use Chunked Search for anything over 5 pages.
  2. Keep documents focused -- one topic per document produces better retrieval results than one large document covering everything.
  3. Test retrieval quality -- use the /search API endpoint to test queries against your documents before deploying to production agents.
  4. Update documents by re-uploading -- to update a document, delete the old version and upload the new one. Documents are immutable after processing.
  5. Use descriptive filenames -- filenames are included in retrieval results, helping the LLM identify which source it is citing.
  6. Use the Files API for automation -- if you are uploading documents programmatically, use POST /api/v1/files instead of the 3-step presigned URL flow.

On this page