Tracing
Integrate Langfuse for tracing LLM calls, tool invocations, and pipeline performance across every voice AI call.
What you'll learn
How to set up Langfuse tracing for your zoxaAI organization, what gets traced during each call (LLM calls, tool calls, TTS, STT), how to view traces, and how multi-organization routing works.
Tracing
zoxaAI integrates with Langfuse for tracing LLM calls, tool calls, and pipeline execution. Traces give you deep visibility into what happens during each call -- every LLM request, every tool invocation, and the latency of each pipeline stage.
Why Tracing Matters
Voice AI pipelines involve multiple components running in real time: speech-to-text, LLM inference, tool execution, and text-to-speech. When something goes wrong -- slow responses, incorrect answers, failed tool calls -- you need visibility into exactly what happened and where the bottleneck occurred.
Tracing provides:
- LLM call inspection -- see the exact prompts, completions, token usage, and latency for every LLM call.
- Tool debugging -- view tool invocation inputs, outputs, and duration.
- Latency analysis -- identify which pipeline stage is causing delays.
- Cost tracking -- understand per-call LLM and provider costs.
- Quality monitoring -- review conversation patterns across many calls.
How It Works
When tracing is enabled, zoxaAI exports OpenTelemetry spans to Langfuse for every pipeline run. Each span captures a discrete operation within the call pipeline. Spans are nested into a trace tree that represents the complete execution timeline for a single call.
Trace: Pipeline Run (call_a1b2c3d4)
|-- Span: LLM Generation (GPT-4.1, 350ms)
|-- Span: Tool Call (check_appointment, 120ms)
|-- Span: LLM Generation (GPT-4.1, 280ms)
|-- Span: TTS Synthesis (ElevenLabs Flash v2.5, 45ms)
|-- Span: STT Transcription (Soniox stt-rt-v5, 30ms)Setting Up Tracing
Get Your Langfuse Credentials
If you use Langfuse Cloud:
- Go to cloud.langfuse.com.
- Select your project (or create one).
- Navigate to Settings > API Keys.
- Copy the Public Key and Secret Key.
If you self-host Langfuse, use your instance URL as the Host.
Open Organization Settings
In the zoxaAI dashboard, navigate to Settings in the sidebar.
Configure Langfuse Credentials
Find the Langfuse section and enter your credentials:
| Field | Description | Example |
|---|---|---|
| Host | Your Langfuse instance URL | https://cloud.langfuse.com |
| Public Key | Your Langfuse project's public key | pk-lf-... |
| Secret Key | Your Langfuse project's secret key | sk-lf-... |
Save
Click Save. All calls from your organization will now be traced to your Langfuse project.
Tracing starts immediately
Once credentials are saved, tracing begins on the next call. Existing completed calls are not retroactively traced.
What Gets Traced
Each pipeline run generates spans for every major component:
| Component | Span Type | Details Captured |
|---|---|---|
| LLM | Generation | Model name, prompt/completion text, tokens in/out, latency, cost |
| Tools / Functions | Span | Tool name, input arguments, output result, duration |
| TTS | Span | Provider, voice name, characters synthesized, latency |
| STT | Span | Provider, model, transcription result, latency |
| Pipeline | Trace (root) | Overall call duration, frame count, error status |
Every span is tagged with the organization ID (zoxa.org_id) for filtering in Langfuse.
Viewing Traces
After tracing is configured, each call's detail page in zoxaAI includes a Trace URL that links directly to the trace in your Langfuse dashboard.
In Langfuse
Once in Langfuse, you can:
| Action | Description |
|---|---|
| View span tree | See the complete nested execution timeline for a call |
| Inspect LLM calls | Read the exact prompts, system messages, and completions |
| Analyze token usage | See prompt tokens, completion tokens, and estimated cost per LLM call |
| Measure latency | Identify slow spans (e.g., a tool call that took 2 seconds) |
| Filter by metadata | Filter traces by organization ID, model, or custom tags |
| Compare traces | Compare execution patterns across different calls |
| Track costs | Aggregate LLM costs across all calls |
Multi-Organization Routing
zoxaAI supports per-organization Langfuse credentials. Each organization's spans are routed to their own Langfuse project independently.
| Scenario | Behavior |
|---|---|
| Organization A has Langfuse credentials configured | Organization A's call traces go to Organization A's Langfuse project |
| Organization B has Langfuse credentials configured | Organization B's call traces go to Organization B's Langfuse project |
| Organization C has no Langfuse credentials | Organization C's traces fall back to the platform-level default exporter (if one exists) |
There is no cross-organization data leakage -- each organization's traces are completely isolated.
Platform-Level Configuration
Platform administrators can configure a default Langfuse exporter via environment variables. These serve as the fallback when an organization does not have its own credentials configured.
| Environment Variable | Description | Default |
|---|---|---|
ENABLE_TRACING | Set to true to enable the tracing subsystem | false |
LANGFUSE_HOST | Default Langfuse host URL | -- |
LANGFUSE_PUBLIC_KEY | Default Langfuse public key | -- |
LANGFUSE_SECRET_KEY | Default Langfuse secret key | -- |
Platform-level vs. organization-level
Organization-level credentials always take priority over platform-level environment variables. The environment variables are only used as a fallback for organizations that have not configured their own Langfuse project.
Use Cases
Debugging Slow Responses
If users report that the agent takes too long to respond:
- Open the call in Call History and click the Trace URL.
- Look at the span durations for LLM, TTS, and STT.
- Identify which component has the highest latency.
- If the LLM is slow, consider switching to a faster model.
- If TTS is slow, try a different TTS provider or voice.
Investigating Wrong Answers
If the agent gives incorrect answers:
- Open the trace for the problematic call.
- Inspect the LLM Generation span to see the full prompt and completion.
- Check if the system prompt was correct and if the right context was provided.
- Review tool call spans to see if tools returned the expected data.
Cost Optimization
- In Langfuse, filter traces by date range and aggregate token usage.
- Identify which calls consume the most tokens.
- Consider optimizing system prompts to reduce token count.
- Evaluate whether a smaller/cheaper model can handle your use case.
Troubleshooting
| Issue | Cause | Solution |
|---|---|---|
| No traces appearing in Langfuse | Credentials not saved or incorrect | Verify the Host, Public Key, and Secret Key in Settings |
| Traces appear but are incomplete | Call ended abnormally (error/crash) | Check the call's ended reason in Call History |
| "Trace URL" not shown on call detail | Tracing not configured for the organization | Add Langfuse credentials in Settings |
| Traces going to wrong project | Credentials point to a different Langfuse project | Update the Public Key and Secret Key to match the correct project |