Browser flow — end-to-end
Complete recipe — create a WebRTC call, complete the SmallWebRTC SDP handshake from the browser, react to lifecycle events.
Voice agent embedded in a web page. Your backend creates the call and returns the offerUrl to the browser; the browser completes the SmallWebRTC SDP handshake against it and audio starts flowing.
This is the right path for: in-product voice assistants, marketing-page demos, dashboard "talk to your data" widgets — anywhere the user is already on your site.
Backend — create the call
The browser should not see your X-API-Key. Do this server-side and return only the callId + offerUrl to the browser.
// server.ts — runs in your Node backend
app.post("/api/voice/start", async (req, res) => {
const r = await fetch("https://dashboard.zoxa.ai/api/v1/call", {
method: "POST",
headers: {
"X-API-Key": process.env.ZOXA_API_KEY, // backend-only secret
"Content-Type": "application/json",
},
body: JSON.stringify({
transport: "webrtc",
agentId: "<agent.uuid>",
contextVariables: {
userId: req.user.id.toString(),
currentPage: req.body.page ?? "unknown",
},
}),
});
const call = await r.json();
// Hand only the signaling fields to the browser:
res.json({ callId: call.callId, offerUrl: call.offerUrl });
});Browser — complete the SmallWebRTC handshake
Create an RTCPeerConnection, add the mic track, and POST the SDP offer to offerUrl. The pipeline launches the moment the offer arrives.
Proxy the signaling through your backend
offerUrl is a relative path (/api/v1/webrtc/offer/{callId}) on the zoxaAI API, and it requires your X-API-Key. Either proxy the POST/PATCH through your backend, or issue a short-lived scoped token — never ship your raw API key to the browser.
<script type="module">
const startBtn = document.querySelector("#talk");
const endBtn = document.querySelector("#stop");
let pc;
startBtn.addEventListener("click", async () => {
const { callId, offerUrl } = await fetch("/api/voice/start", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ page: location.pathname }),
}).then((r) => r.json());
pc = new RTCPeerConnection();
// Play the agent's audio.
pc.ontrack = ({ streams: [stream] }) => {
document.querySelector("#agent-audio").srcObject = stream;
};
// Send the user's mic.
const mic = await navigator.mediaDevices.getUserMedia({ audio: true });
mic.getTracks().forEach((t) => pc.addTrack(t, mic));
// Offer → answer. Proxy `offerUrl` through your backend (it needs the API key).
const offer = await pc.createOffer();
await pc.setLocalDescription(offer);
const answer = await fetch(`/api/voice/signal?path=${encodeURIComponent(offerUrl)}`, {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ sdp: offer.sdp, type: offer.type }),
}).then((r) => r.json());
await pc.setRemoteDescription(answer);
// Trickle ICE candidates (PATCH the same offerUrl).
pc.onicecandidate = ({ candidate }) => {
if (!candidate) return;
fetch(`/api/voice/signal?path=${encodeURIComponent(offerUrl)}`, {
method: "PATCH",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ candidate }),
});
};
startBtn.disabled = true;
endBtn.disabled = false;
});
endBtn.addEventListener("click", () => {
pc?.close();
startBtn.disabled = false;
endBtn.disabled = true;
});
</script>
<audio id="agent-audio" autoplay></audio>The agent's greeting line plays as soon as the peer connection is established (if greeting.speakFirst: "agent").
React to call end on the backend
Set webhook.url on the agent and you get call.ended server-side when the call terminates. Or poll GET /calls/{callId} if the user clicked "End" in your UI:
endBtn.addEventListener("click", async () => {
pc?.close();
await fetch(`/api/voice/${callId}/log-feedback`, { method: "POST" });
});(Your server hits GET /calls/{callId} to read transcript / cost / analysis and stores feedback alongside.)
What zoxaAI handles for you
| Concern | How |
|---|---|
| WebRTC signaling | pipecat's SmallWebRTCRequestHandler answers your offer and manages ICE trickle + renegotiation. |
| NAT traversal | STUN/TURN via the ICE servers the offer endpoint advertises. |
| Browser compatibility | Standard RTCPeerConnection — works on all evergreen browsers. |
| Auth on the WebRTC leg | The offer/PATCH endpoints use your X-API-Key — proxy them server-side. |
| Call recording | Recorded server-side; recordings[].url after the call ends. |
What's specific to WebRTC
| Constraint | Value |
|---|---|
| Max duration | Default 610 seconds. Override via the agent's maxCallDurationS. |
| Signaling | One SDP offer per call (POST), plus trickle-ICE (PATCH). A duplicate offer returns 409. |
Common pitfalls
| Symptom | Fix |
|---|---|
| User connects, agent doesn't speak | The agent waits greeting.userTimeoutS for the user. Set greeting.speakFirst: "agent" for the bot to greet first. |
409 on the offer POST | You POSTed a second offer for a call whose pipeline is already starting — the handshake happens once per call. |
| No agent audio | Make sure you wired pc.ontrack to an <audio autoplay> element before setting the remote description. |
| Audio cuts off early | You hit maxCallDurationS (default 610s). Increase it via the agent config. |
Related
- Calls — WebRTC — endpoint reference and the SmallWebRTC signaling flow
- Choosing a transport — when WebRTC is the right pick
- Example agents — in-product voice helper — copy-pasteable browser-flow agent