zoxaAI
Homepage
API ReferenceExamples

Browser flow — end-to-end

Complete recipe — create a WebRTC call, complete the SmallWebRTC SDP handshake from the browser, react to lifecycle events.

Voice agent embedded in a web page. Your backend creates the call and returns the offerUrl to the browser; the browser completes the SmallWebRTC SDP handshake against it and audio starts flowing.

This is the right path for: in-product voice assistants, marketing-page demos, dashboard "talk to your data" widgets — anywhere the user is already on your site.

Backend — create the call

The browser should not see your X-API-Key. Do this server-side and return only the callId + offerUrl to the browser.

// server.ts — runs in your Node backend
app.post("/api/voice/start", async (req, res) => {
  const r = await fetch("https://dashboard.zoxa.ai/api/v1/call", {
    method: "POST",
    headers: {
      "X-API-Key":    process.env.ZOXA_API_KEY,    // backend-only secret
      "Content-Type": "application/json",
    },
    body: JSON.stringify({
      transport: "webrtc",
      agentId:   "<agent.uuid>",
      contextVariables: {
        userId:      req.user.id.toString(),
        currentPage: req.body.page ?? "unknown",
      },
    }),
  });
  const call = await r.json();
  // Hand only the signaling fields to the browser:
  res.json({ callId: call.callId, offerUrl: call.offerUrl });
});

Browser — complete the SmallWebRTC handshake

Create an RTCPeerConnection, add the mic track, and POST the SDP offer to offerUrl. The pipeline launches the moment the offer arrives.

Proxy the signaling through your backend

offerUrl is a relative path (/api/v1/webrtc/offer/{callId}) on the zoxaAI API, and it requires your X-API-Key. Either proxy the POST/PATCH through your backend, or issue a short-lived scoped token — never ship your raw API key to the browser.

<script type="module">
  const startBtn = document.querySelector("#talk");
  const endBtn   = document.querySelector("#stop");

  let pc;

  startBtn.addEventListener("click", async () => {
    const { callId, offerUrl } = await fetch("/api/voice/start", {
      method: "POST",
      headers: { "Content-Type": "application/json" },
      body: JSON.stringify({ page: location.pathname }),
    }).then((r) => r.json());

    pc = new RTCPeerConnection();

    // Play the agent's audio.
    pc.ontrack = ({ streams: [stream] }) => {
      document.querySelector("#agent-audio").srcObject = stream;
    };

    // Send the user's mic.
    const mic = await navigator.mediaDevices.getUserMedia({ audio: true });
    mic.getTracks().forEach((t) => pc.addTrack(t, mic));

    // Offer → answer. Proxy `offerUrl` through your backend (it needs the API key).
    const offer = await pc.createOffer();
    await pc.setLocalDescription(offer);
    const answer = await fetch(`/api/voice/signal?path=${encodeURIComponent(offerUrl)}`, {
      method: "POST",
      headers: { "Content-Type": "application/json" },
      body: JSON.stringify({ sdp: offer.sdp, type: offer.type }),
    }).then((r) => r.json());
    await pc.setRemoteDescription(answer);

    // Trickle ICE candidates (PATCH the same offerUrl).
    pc.onicecandidate = ({ candidate }) => {
      if (!candidate) return;
      fetch(`/api/voice/signal?path=${encodeURIComponent(offerUrl)}`, {
        method: "PATCH",
        headers: { "Content-Type": "application/json" },
        body: JSON.stringify({ candidate }),
      });
    };

    startBtn.disabled = true;
    endBtn.disabled = false;
  });

  endBtn.addEventListener("click", () => {
    pc?.close();
    startBtn.disabled = false;
    endBtn.disabled = true;
  });
</script>
<audio id="agent-audio" autoplay></audio>

The agent's greeting line plays as soon as the peer connection is established (if greeting.speakFirst: "agent").

React to call end on the backend

Set webhook.url on the agent and you get call.ended server-side when the call terminates. Or poll GET /calls/{callId} if the user clicked "End" in your UI:

endBtn.addEventListener("click", async () => {
  pc?.close();
  await fetch(`/api/voice/${callId}/log-feedback`, { method: "POST" });
});

(Your server hits GET /calls/{callId} to read transcript / cost / analysis and stores feedback alongside.)

What zoxaAI handles for you

ConcernHow
WebRTC signalingpipecat's SmallWebRTCRequestHandler answers your offer and manages ICE trickle + renegotiation.
NAT traversalSTUN/TURN via the ICE servers the offer endpoint advertises.
Browser compatibilityStandard RTCPeerConnection — works on all evergreen browsers.
Auth on the WebRTC legThe offer/PATCH endpoints use your X-API-Key — proxy them server-side.
Call recordingRecorded server-side; recordings[].url after the call ends.

What's specific to WebRTC

ConstraintValue
Max durationDefault 610 seconds. Override via the agent's maxCallDurationS.
SignalingOne SDP offer per call (POST), plus trickle-ICE (PATCH). A duplicate offer returns 409.

Common pitfalls

SymptomFix
User connects, agent doesn't speakThe agent waits greeting.userTimeoutS for the user. Set greeting.speakFirst: "agent" for the bot to greet first.
409 on the offer POSTYou POSTed a second offer for a call whose pipeline is already starting — the handshake happens once per call.
No agent audioMake sure you wired pc.ontrack to an <audio autoplay> element before setting the remote description.
Audio cuts off earlyYou hit maxCallDurationS (default 610s). Increase it via the agent config.

On this page