LLM Providers
Every language model provider supported by zoxaAI -- available models, default configuration, and provider-specific options.
Overview
zoxaAI provides managed access to 4 LLM providers, each with a curated catalog of production-tested models. Each agent is configured with exactly one LLM provider and model.
Curated catalog — no free-text model ids
Model selection is a curated catalog, not free text. You pick from the models listed on each provider below; arbitrary or newly-released model ids that aren't in the catalog are rejected at save time with a 422 pointing at GET /api/v1/user/configurations/providers. A provider only appears in the picker when it has at least one active model in the catalog.
Every model in the catalog carries metadata — latency (time-to-first-token / time-to-first-byte), throughput, an intelligence index, a tool-calling score, word error rate, context window, and per-unit cost — stored alongside its pricing. The catalog endpoint returns this metadata so you can compare models on speed, cost, and quality before choosing.
Provider Comparison
| Provider | Default Model | Notes |
|---|---|---|
| OpenAI | gpt-4.1 | Priority service tier, widest model selection |
gemini-2.5-flash | Gemini models including 3.x previews | |
| Qwen | qwen-flash | Alibaba DashScope; strong tool calling, includes MoE agentic variants |
| Anthropic | claude-sonnet-4-6 | Claude model family |
Shared Configuration Fields
Every LLM provider supports these common fields:
| Field | Type | Default | Description |
|---|---|---|---|
model | string | Per-provider | The model identifier. Must be one of the models in the provider's catalog (below). |
temperature | float | 1.0 | Sampling temperature. Lower values produce more deterministic output. Range: 0.0 -- 2.0 for OpenAI, Google, and Qwen. Anthropic clamps to 0.0 -- 1.0 (see the Anthropic section below). |
max_tokens | integer | 251 | Maximum number of tokens in the completion. A spoken turn is a sentence or two (~40 tokens), so the default leaves generous headroom while still stopping a runaway monologue. Range: 1 -- 4096. |
prewarm | bool | true | Sends one tiny out-of-band request at pipeline start so the first real turn skips the provider's cold start (TLS handshake + prompt prefill / routing). Default on for a large turn-1 latency win. |
OpenAI
Default model: gpt-4.1
| Model | Notes |
|---|---|
gpt-5.4 | Mid-tier flagship — billed at Priority ($5 in / $30 out per 1M) |
gpt-5.4-mini | Voice-friendly, billed at Priority ($1.50 in / $9 out per 1M) |
gpt-5.4-nano | Cheapest current-gen, sub-second TTFT (Standard tier) |
gpt-6-luna | GPT-6 family, high-volume and lowest cost ($0.10 in / $0.50 out per 1M, Standard tier) — reasoning set to none automatically |
gpt-5.2 | Late-2025 family — billed at Priority ($3.50 in / $28 out per 1M) |
gpt-5.1 | Late-2025, configurable reasoning effort — billed at Priority ($2.50 in / $20 out per 1M) |
gpt-4.1 | Default -- reliable, well-tested for voice — billed at Priority ($3.50 in / $14 out per 1M) |
gpt-4.1-mini | Cost-efficient, lower latency — billed at Priority ($0.70 in / $2.80 out per 1M) |
Priority service tier
Every OpenAI model except gpt-5.4-nano and gpt-6-luna is dispatched with service_tier: priority (OpenAI's Fast tier) for lower and more consistent voice TTFT. OpenAI charges more for Priority calls; the cost calculator already reflects the Priority rates so the dashboard shows what OpenAI actually invoices. gpt-5.4-nano and gpt-6-luna bill — and run — at Standard.
Parameter shapes per family
The factory routes parameters per OpenAI's /v1/chat/completions constraints — picking the wrong shape returns 400:
-
gpt-6 family (
gpt-6-luna) — function tools work on chat-completions only withreasoning_effort: none, which the factory always sends (it is also the right setting for voice). Customtemperature+max_completion_tokens. -
gpt-5.x point releases (
gpt-5.1/5.2/5.4/5.4-mini/5.4-nano) — OpenAI moved the tools-plus-reasoning combo to/v1/responses. On chat-completions these models rejectreasoning_effortwhenever tools are present (which our pipeline always sends). They accept a customtemperature; the factory passesmax_completion_tokens+temperature. -
gpt-4.1 family (
gpt-4.1,gpt-4.1-mini) — classic shape: customtemperature+max_tokens.
Google (Gemini)
Default model: gemini-2.5-flash
| Model | Notes |
|---|---|
gemini-2.5-flash | Default -- balanced speed/quality |
gemini-2.5-flash-lite | Lightweight flash variant |
gemini-3.5-flash-lite | Low-latency Gemini 3.5 lightweight model, thinking kept at minimal |
gemini-3-flash-preview | Next-gen flash preview |
gemini-3.1-flash-lite | Lightweight Gemini 3.1 model |
Auth: Managed by the platform.
Thinking kept minimal for voice
Gemini 2.5 models run with thinking_budget: 0 (thinking off), and Gemini 3.x models run with thinking_level: minimal, so replies start without a reasoning phase.
Qwen (Alibaba)
Default model: qwen-flash
Endpoint: https://dashscope-intl.aliyuncs.com/compatible-mode/v1 (Alibaba Cloud Model Studio, International / Singapore). Qwen exposes an OpenAI-compatible chat-completions API, so streaming and function calling work identically to OpenAI — pipecat's QwenLLMService inherits from OpenAILLMService.
The catalog is voice-curated: every model below has sub-second TTFT and reliable multi-turn tool calling, so the streaming-TTS handoff stays tight on live calls.
| Model | Notes |
|---|---|
qwen-flash | Default -- low-latency, cheapest, full tool calling |
qwen3.6-flash | Latest flash; thinking mode disabled for voice |
qwen3.5-flash | Flat-rate flash variant |
qwen3.6-35b-a3b | Open-weight MoE (35B total / 3B active) — voice-fast |
qwen-plus | Balanced reasoning, still voice-suitable |
qwen3.7-plus | Newest plus-tier variant |
enable_thinking is forced to false for every Qwen call. Every Qwen model in the catalog is a hybrid thinking model: with thinking on, it streams its reasoning before the actual reply, adding 1-2 seconds to TTFT and breaking streaming TTS — voice is hard-incompatible with that mode. With enable_thinking: false every catalog model answers directly. The flag is passed through OpenAI's extra_body channel; it is NOT a direct chat-completions kwarg.
Anthropic
Default model: claude-sonnet-4-6
| Model | Notes |
|---|---|
claude-sonnet-4-6 | Default -- latest Sonnet, balanced quality and speed |
claude-haiku-4-5 | Fastest, most cost-efficient |
Auth: Managed by the platform.
Temperature range is 0.0 -- 1.0. Anthropic's Messages API rejects temperatures above 1.0. When you select an Anthropic model in the agent editor, the Temperature slider caps at 1.0 and any stored value above 1.0 is clamped. If you call the API directly, ensure temperature ≤ 1.0 for any Anthropic provider config.