zoxaAI
Homepage

LLM Providers

Every language model provider supported by zoxaAI -- available models, default configuration, and provider-specific options.

Overview

zoxaAI provides managed access to 4 LLM providers, each with a curated catalog of production-tested models. Each agent is configured with exactly one LLM provider and model.

Curated catalog — no free-text model ids

Model selection is a curated catalog, not free text. You pick from the models listed on each provider below; arbitrary or newly-released model ids that aren't in the catalog are rejected at save time with a 422 pointing at GET /api/v1/user/configurations/providers. A provider only appears in the picker when it has at least one active model in the catalog.

Every model in the catalog carries metadata — latency (time-to-first-token / time-to-first-byte), throughput, an intelligence index, a tool-calling score, word error rate, context window, and per-unit cost — stored alongside its pricing. The catalog endpoint returns this metadata so you can compare models on speed, cost, and quality before choosing.


Provider Comparison

ProviderDefault ModelNotes
OpenAIgpt-4.1Priority service tier, widest model selection
Googlegemini-2.5-flashGemini models including 3.x previews
Qwenqwen-flashAlibaba DashScope; strong tool calling, includes MoE agentic variants
Anthropicclaude-sonnet-4-6Claude model family

Shared Configuration Fields

Every LLM provider supports these common fields:

FieldTypeDefaultDescription
modelstringPer-providerThe model identifier. Must be one of the models in the provider's catalog (below).
temperaturefloat1.0Sampling temperature. Lower values produce more deterministic output. Range: 0.0 -- 2.0 for OpenAI, Google, and Qwen. Anthropic clamps to 0.0 -- 1.0 (see the Anthropic section below).
max_tokensinteger251Maximum number of tokens in the completion. A spoken turn is a sentence or two (~40 tokens), so the default leaves generous headroom while still stopping a runaway monologue. Range: 1 -- 4096.
prewarmbooltrueSends one tiny out-of-band request at pipeline start so the first real turn skips the provider's cold start (TLS handshake + prompt prefill / routing). Default on for a large turn-1 latency win.

OpenAI

Default model: gpt-4.1

ModelNotes
gpt-5.4Mid-tier flagship — billed at Priority ($5 in / $30 out per 1M)
gpt-5.4-miniVoice-friendly, billed at Priority ($1.50 in / $9 out per 1M)
gpt-5.4-nanoCheapest current-gen, sub-second TTFT (Standard tier)
gpt-6-lunaGPT-6 family, high-volume and lowest cost ($0.10 in / $0.50 out per 1M, Standard tier) — reasoning set to none automatically
gpt-5.2Late-2025 family — billed at Priority ($3.50 in / $28 out per 1M)
gpt-5.1Late-2025, configurable reasoning effort — billed at Priority ($2.50 in / $20 out per 1M)
gpt-4.1Default -- reliable, well-tested for voice — billed at Priority ($3.50 in / $14 out per 1M)
gpt-4.1-miniCost-efficient, lower latency — billed at Priority ($0.70 in / $2.80 out per 1M)

Priority service tier

Every OpenAI model except gpt-5.4-nano and gpt-6-luna is dispatched with service_tier: priority (OpenAI's Fast tier) for lower and more consistent voice TTFT. OpenAI charges more for Priority calls; the cost calculator already reflects the Priority rates so the dashboard shows what OpenAI actually invoices. gpt-5.4-nano and gpt-6-luna bill — and run — at Standard.

Parameter shapes per family

The factory routes parameters per OpenAI's /v1/chat/completions constraints — picking the wrong shape returns 400:

  • gpt-6 family (gpt-6-luna) — function tools work on chat-completions only with reasoning_effort: none, which the factory always sends (it is also the right setting for voice). Custom temperature + max_completion_tokens.

  • gpt-5.x point releases (gpt-5.1 / 5.2 / 5.4 / 5.4-mini / 5.4-nano) — OpenAI moved the tools-plus-reasoning combo to /v1/responses. On chat-completions these models reject reasoning_effort whenever tools are present (which our pipeline always sends). They accept a custom temperature; the factory passes max_completion_tokens + temperature.

  • gpt-4.1 family (gpt-4.1, gpt-4.1-mini) — classic shape: custom temperature + max_tokens.


Google (Gemini)

Default model: gemini-2.5-flash

ModelNotes
gemini-2.5-flashDefault -- balanced speed/quality
gemini-2.5-flash-liteLightweight flash variant
gemini-3.5-flash-liteLow-latency Gemini 3.5 lightweight model, thinking kept at minimal
gemini-3-flash-previewNext-gen flash preview
gemini-3.1-flash-liteLightweight Gemini 3.1 model

Auth: Managed by the platform.

Thinking kept minimal for voice

Gemini 2.5 models run with thinking_budget: 0 (thinking off), and Gemini 3.x models run with thinking_level: minimal, so replies start without a reasoning phase.


Qwen (Alibaba)

Default model: qwen-flash

Endpoint: https://dashscope-intl.aliyuncs.com/compatible-mode/v1 (Alibaba Cloud Model Studio, International / Singapore). Qwen exposes an OpenAI-compatible chat-completions API, so streaming and function calling work identically to OpenAI — pipecat's QwenLLMService inherits from OpenAILLMService.

The catalog is voice-curated: every model below has sub-second TTFT and reliable multi-turn tool calling, so the streaming-TTS handoff stays tight on live calls.

ModelNotes
qwen-flashDefault -- low-latency, cheapest, full tool calling
qwen3.6-flashLatest flash; thinking mode disabled for voice
qwen3.5-flashFlat-rate flash variant
qwen3.6-35b-a3bOpen-weight MoE (35B total / 3B active) — voice-fast
qwen-plusBalanced reasoning, still voice-suitable
qwen3.7-plusNewest plus-tier variant

enable_thinking is forced to false for every Qwen call. Every Qwen model in the catalog is a hybrid thinking model: with thinking on, it streams its reasoning before the actual reply, adding 1-2 seconds to TTFT and breaking streaming TTS — voice is hard-incompatible with that mode. With enable_thinking: false every catalog model answers directly. The flag is passed through OpenAI's extra_body channel; it is NOT a direct chat-completions kwarg.


Anthropic

Default model: claude-sonnet-4-6

ModelNotes
claude-sonnet-4-6Default -- latest Sonnet, balanced quality and speed
claude-haiku-4-5Fastest, most cost-efficient

Auth: Managed by the platform.

Temperature range is 0.0 -- 1.0. Anthropic's Messages API rejects temperatures above 1.0. When you select an Anthropic model in the agent editor, the Temperature slider caps at 1.0 and any stored value above 1.0 is clamped. If you call the API directly, ensure temperature ≤ 1.0 for any Anthropic provider config.


Choosing a Provider

On this page