WebSocket audio protocol
WS /api/v1/ws/audio/{callId} — bidirectional raw-PCM audio + JSON events. Auth via api_key query string; close codes for invalid call states.
WS /api/v1/ws/audio/{callId}Bidirectional audio stream for a call created with transport=websocket. Connect with the callId returned from POST /call, and raw PCM flows immediately in both directions. No SDP, no ICE.
Connection URL
wss://dashboard.zoxa.ai/api/v1/ws/audio/{callId}?api_key=zsk_...| Component | Description |
|---|---|
callId | The id returned from POST /call (the one with audioUrl). |
api_key (query) | Browser WebSocket handshakes can't carry custom headers, so the key goes in the query string. Server-side clients can use the X-API-Key header instead. |
Audio format
The wire format is fixed:
| Direction | Format |
|---|---|
| Client → Server | pcm_s16le, raw, 16 000 Hz, mono |
| Server → Client | pcm_s16le, raw, 24 000 Hz, mono |
Every binary frame you send must be 16 kHz PCM-16. Every binary frame the server sends back is 24 kHz PCM-16 — the agent's voice at full synthesis quality; play it out at 24 kHz (or resample on your side if you're bridging to a narrowband system).
Frame types
| Direction | Frame type | Payload | Purpose |
|---|---|---|---|
| Client → Server | binary | raw PCM-16 at 16 kHz | Microphone audio. Send small frames (~20 ms) for low latency. |
| Server → Client | binary | raw PCM-16 at 24 kHz | Agent speech audio. Play out as-is at 24 kHz. |
| Server → Client | text | JSON event (see below) | Live transcription, agent text, errors. |
Server → Client JSON events
All text frames are JSON. The type field is the discriminator.
rtf-user-transcription
Streaming transcription of the user's speech.
{
"type": "rtf-user-transcription",
"payload": {
"text": "hello, I wanted to ask about",
"final": false
}
}| Field | Type | Description |
|---|---|---|
payload.text | string | The transcript so far for the current utterance. |
payload.final | boolean | false for interim chunks, true when the STT marks the utterance complete. Use final=true to commit a turn to your UI. |
rtf-bot-text
Agent's reply text, streamed alongside the audio so you can render a transcript live.
{
"type": "rtf-bot-text",
"payload": { "text": "Sure — what would you like to ask?" }
}error
A non-fatal error happened during pipeline execution. The connection may still be alive.
{ "type": "error", "message": "Call failed" }If the connection is unrecoverable, the server will follow with a WS close.
WebSocket close codes
| Code | Reason text | When |
|---|---|---|
1000 | (clean) | Call ended normally — agent hung up, user disconnected gracefully, or duration limit hit. |
1008 | "Invalid call" | callId doesn't exist in your organization. |
1008 | "Call not available" | Call is not in pending status (already active in another WS, already completed, etc.). |
1008 | "Call transport mismatch" | The call was created with transport=webrtc, not websocket. |
A 1008 close means the connection is rejected at handshake — your onopen may not fire.
Lifecycle on connect
- Server validates the
callIdis for your org and is inpendingstatus. - Server
accept()s the WS. - Server marks the call
activeand starts the pipeline. - Audio + text events flow until either side closes.
- If the pipeline never started (an exception before the runner), the server marks the call
connectionStatus=cancelled(withendedReasonleftnull— the call never connected) in the finally block so no row gets stuck inpending.
Duration limits
There is no fixed WS duration cap at the endpoint level — the call runs as long as you keep the WS open and the pipeline stays healthy. The agent config's maxCallDurationS does cap it from above: when the pipeline hits that limit, the agent ends the call gracefully and the server closes the WS with code 1000.
To extend, raise maxCallDurationS on the agent (POST /agents) or on the inline agent sent with POST /call.
Example: minimal Node.js client
import WebSocket from "ws";
const ws = new WebSocket(
`wss://dashboard.zoxa.ai/api/v1/ws/audio/${callId}?api_key=zsk_...`,
);
ws.on("open", () => {
// Start sending mic PCM — Buffer of int16 little-endian samples at 16 kHz
micStream.on("data", (chunk) => ws.send(chunk));
});
ws.on("message", (data, isBinary) => {
if (isBinary) {
speaker.write(data); // play agent audio (24 kHz PCM)
} else {
const event = JSON.parse(data.toString());
if (event.type === "rtf-user-transcription") {
console.log("you:", event.payload.text, event.payload.final ? "(final)" : "");
} else if (event.type === "rtf-bot-text") {
console.log("bot:", event.payload.text);
} else if (event.type === "error") {
console.error("call error:", event.message);
}
}
});
ws.on("close", (code, reason) => {
console.log("call ended", code, reason.toString());
});Example: minimal Python client
import asyncio, json
import websockets
async def main(call_id: str):
url = f"wss://dashboard.zoxa.ai/api/v1/ws/audio/{call_id}?api_key=zsk_..."
async with websockets.connect(url) as ws:
async def send_mic():
async for chunk in mic_pcm_chunks(): # int16 LE 16 kHz frames
await ws.send(chunk)
async def recv_loop():
async for msg in ws:
if isinstance(msg, bytes):
play_pcm(msg) # agent audio (24 kHz PCM)
else:
event = json.loads(msg)
if event["type"] == "rtf-user-transcription":
print("you:", event["payload"]["text"])
elif event["type"] == "rtf-bot-text":
print("bot:", event["payload"]["text"])
await asyncio.gather(send_mic(), recv_loop())
asyncio.run(main("call_..."))Tips
| Tip | Why |
|---|---|
| Send small audio frames (~20 ms = 640 bytes at pcm_s16le/16 kHz) | Latency budget — bigger frames stall the VAD/STT pipeline. |
| Don't buffer received audio before playing | The server has already paced it for real-time playback. |
Treat final: true as a commit point | The STT may revise interim transcriptions; final is the durable text. |
Close cleanly with code 1000 when done | Avoids the cancelled-cleanup path on the server. |
Related
POST /call(websocket transport) — create the call before connectingGET /calls/{callId}— full transcript / recording / latency after the call- WebRTC mode — browser-friendly alternative