zoxaAI
Homepage

Knowledge Base Tool

Configure tools that search your uploaded documents during voice calls, giving your agent instant access to product docs, policies, FAQs, and any other reference material via RAG retrieval.

What you'll learn

  • How to create a Knowledge Base tool in the dashboard and attach it to agents
  • How to use inline query tools via the API for transient and override calls
  • How multi-group KB routing works when you have multiple document sets
  • The full retrieval flow that runs during a live call
  • Validation rules and best practices

The Knowledge Base tool lets your agent search uploaded documents during a live call. When the caller asks a question that requires specific information -- a return policy, product spec, pricing detail -- the LLM calls the retrieval function, the platform searches your documents using RAG, and the relevant content is returned so the agent can answer accurately.

There are two ways to use knowledge base retrieval during calls:

  1. Dashboard tool -- create a persistent knowledge_base tool in the Tools page and attach it to agents, just like any other tool.
  2. Inline query tool (API) -- pass a {"type": "query", ...} tool inline when making a call via the API, useful for transient or override agent configurations.

Creating via Dashboard

Use the dashboard when you want a reusable, persistent tool that can be attached to multiple agents.

Navigate to Tools in the dashboard sidebar and click Create Tool.

Select Knowledge Base as the tool type.

Enter a name and description. The description tells the LLM when to search the knowledge base -- be specific about the kind of questions this tool answers.

In the Knowledge Base Configuration section, select the documents the agent should search. Only documents with status completed are available for selection.

Optionally add up to 5 Filler Messages — short lines like "Let me check that for you..." spoken while the search runs, so the caller never hears dead air. They play in order, one per lookup, then cycle, so back-to-back searches in the same call don't sound identical. Leave empty for silence.

Configure retrieval settings (number of results, similarity thresholds) as needed.

Save the tool.

Once saved, attach the tool to any agent. During calls, the LLM sees a retrieve_from_knowledge_base function and calls it when the conversation matches the tool's description.

Documents must be uploaded and fully processed before they can be selected. See the Knowledge Base page for how to upload and process documents.


Using via API (Inline Query Tools)

When making calls via the API with a transient or override agent configuration, you can pass knowledge base tools inline -- no need to pre-create them in the dashboard.

Inline Query Tool Schema

{
  "type": "query",
  "function": {
    "name": "search_product_docs",
    "description": "Search product documentation for answers to customer questions"
  },
  "knowledgeBases": [
    {
      "name": "Product Docs",
      "description": "Product documentation and user guides",
      "documentIds": ["uuid-1", "uuid-2"]
    }
  ],
  "timeout": 5
}

Field Reference

FieldTypeRequiredDescription
type"query"YesIdentifies this as a knowledge base query tool.
function.namestringYesThe function name the LLM sees. Must be 1-64 characters, alphanumeric with underscores and dashes.
function.descriptionstringYesTells the LLM when to use this tool.
knowledgeBasesarrayYesOne or more knowledge base groups to search.
knowledgeBases[].namestringYesDisplay name for the KB group.
knowledgeBases[].descriptionstringYesDescribes the content in this group. Used by the KB router when multiple groups are defined.
knowledgeBases[].documentIdsstring[]YesArray of document UUIDs. These are the file IDs returned from POST /api/v1/files or the document_uuid from the Knowledge Base API.
messagesstring[]NoFiller lines spoken while retrieval runs. Sequential — the first lookup in a call speaks the first entry, the next lookup the second, wrapping around. Max 5 entries; [] (default) is silent.
timeoutintegerNoTimeout in seconds (1-30). Default: 5.

Example: Call with Inline KB Tool

curl -X POST https://dashboard.zoxa.ai/api/v1/call \
  -H "X-API-Key: YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "phoneNumber": "+15551234567",
    "agent": {
      "name": "Support Agent",
      "systemPrompt": "You are a helpful support agent. Answer questions using the product documentation.",
      "llm": { "provider": "openai", "model": "gpt-5.4-mini" },
      "tts": { "provider": "elevenlabs", "model": "eleven_flash_v2_5" },
      "stt": { "provider": "soniox", "model": "stt-rt-v5" },
      "tools": [
        {
          "type": "query",
          "function": {
            "name": "search_product_docs",
            "description": "Search product documentation when the caller asks about features, setup, troubleshooting, or usage."
          },
          "knowledgeBases": [
            {
              "name": "Product Docs",
              "description": "Product documentation, setup guides, and troubleshooting articles",
              "documentIds": ["a1b2c3d4-e5f6-7890-abcd-ef1234567890"]
            }
          ],
          "timeout": 5
        },
        { "type": "endCall" }
      ]
    }
  }'

All Inline Tool Types

The tools array in a transient or override agent config accepts these formats:

FormatDescriptionExample
UUID stringReference a pre-created tool by its UUID"a1b2c3d4-..."
{"type": "function", ...}Inline HTTP API tool (calls an external URL)See HTTP API Tool
{"type": "query", ...}Inline KB query tool (retrieval handled internally)See schema above
{"type": "endCall"}Inline end call tool{"type": "endCall"}

Inline query tools are handled entirely by the platform -- no external HTTP call is made. The platform embeds the query, searches the specified documents, and returns results to the LLM. This is different from inline function tools, which make HTTP requests to a server.url you provide.


Multi-Group KB Routing

When you define multiple knowledgeBases entries in a single query tool, the platform uses an intelligent routing step to determine which group(s) to search.

How It Works

Caller: "What's the return policy for electronics?"
    |
    v
LLM calls search tool with query: "return policy for electronics"
    |
    v
KB Router (gpt-4.1-mini) classifies the query:
  - Group 1: "Return Policies" --> MATCH
  - Group 2: "Product Specs"   --> NO MATCH
    |
    v
Platform searches only the "Return Policies" documents
    |
    v
Results returned to LLM --> agent answers the caller

Example: Multiple KB Groups

{
  "type": "query",
  "function": {
    "name": "search_company_docs",
    "description": "Search company documentation for policies, product info, or HR questions"
  },
  "knowledgeBases": [
    {
      "name": "Product Catalog",
      "description": "Product specifications, features, pricing, and availability",
      "documentIds": ["uuid-products-1", "uuid-products-2"]
    },
    {
      "name": "Company Policies",
      "description": "Return policies, warranty terms, shipping policies, and customer service guidelines",
      "documentIds": ["uuid-policies-1"]
    },
    {
      "name": "HR Handbook",
      "description": "Employee benefits, PTO policies, onboarding procedures, and company values",
      "documentIds": ["uuid-hr-1", "uuid-hr-2", "uuid-hr-3"]
    }
  ],
  "timeout": 8
}

When the caller asks "What's the PTO policy?", the KB router identifies the "HR Handbook" group as the best match and searches only those documents, improving both speed and relevance.

Multi-group routing adds a small amount of latency because the router must classify the query before retrieval begins. If you only have one logical document set, use a single knowledgeBases entry to skip the routing step entirely.


How Retrieval Works During a Call

Regardless of whether you use a dashboard tool or an inline query tool, the retrieval flow during a live call follows the same steps:

LLM triggers retrieval -- the LLM calls the registered retrieval function (e.g., retrieve_from_knowledge_base or your custom function name) with a search query it formulates from the conversation.

KB routing (if multiple groups) -- if multiple knowledgeBases groups are defined, a router model (gpt-4.1-mini) classifies the query to determine which group(s) to search.

Document search -- the platform searches the selected documents based on their retrieval mode:

  • Chunked documents: the query is embedded using text-embedding-3-small, and vector similarity search returns the top matching chunks.
  • Full Document mode documents: the complete document text is returned with a similarity score of 1.0.

Results returned to LLM -- the retrieved text (with source filenames and similarity scores) is passed back to the LLM as the tool result.

LLM responds -- the LLM incorporates the retrieved information into its spoken response to the caller.


Validation Rules

The platform validates all tools -- including inline query tools -- before a call starts. Invalid tools cause the call to fail with a descriptive error.

Tool Name Rules

RuleDetail
Length1-64 characters
CharactersAlphanumeric, underscores (_), and dashes (-) only
UniquenessNo duplicate tool names within the same agent or call

Reserved Names

These function names are reserved by the platform and cannot be used for custom tools:

  • end_call
  • retrieve_from_knowledge_base
  • safe_calculator

Document Validation

RuleDetail
ExistenceAll documentIds must reference documents that exist in your organization
StatusAll referenced documents must have processing status completed

Timeout Limits

Tool TypeRangeUnit
Inline function tools (type: "function")1000 - 30000Milliseconds
Inline query tools (type: "query")1 - 30Seconds

URL Security (Function Tools)

Server URLs for inline function tools cannot point to private or internal IP addresses (localhost, 10.x.x.x, 172.16-31.x.x, 192.168.x.x). This does not apply to query tools, which have no external URL.

HTTP Methods (Function Tools)

Only these HTTP methods are allowed for inline function tools: GET, POST, PUT, PATCH, DELETE.


Best Practices

  1. Write specific tool descriptions -- tell the LLM exactly what kind of questions this tool answers. "Search the knowledge base" is too vague. "Search product documentation for feature details, setup instructions, and troubleshooting steps" is much better.

  2. Use a single KB group when possible -- if all your documents cover the same domain, put them in one group. Multi-group routing adds latency and is only valuable when you have clearly distinct document categories.

  3. Keep documents focused -- one topic per document produces better retrieval results than a single large document covering everything. A 3-page "Return Policy" document retrieves more accurately than a 50-page "All Company Policies" document.

  4. Increase timeout for large document sets -- the default 5-second timeout works for most cases, but if you have many chunked documents or multiple KB groups with routing, consider increasing to 8-10 seconds.

  5. Ensure documents are processed before calling -- the validator checks that all documentIds have status completed. If you just uploaded a document, wait for processing to finish before making calls that reference it.

  6. Test retrieval quality first -- use the Knowledge Base search API to test queries against your documents before deploying to a live agent.

  7. Combine with other tools -- KB tools work alongside HTTP API, Call Transfer, and End Call tools. A common pattern is a KB tool for answering questions plus an End Call tool for wrapping up the conversation.


On this page