zoxaAI
Homepage

Tracing

Integrate Langfuse for tracing LLM calls, tool invocations, and pipeline performance across every voice AI call.

What you'll learn

How to set up Langfuse tracing for your zoxaAI organization, what gets traced during each call (LLM calls, tool calls, TTS, STT), how to view traces, and how multi-organization routing works.

Tracing

zoxaAI integrates with Langfuse for tracing LLM calls, tool calls, and pipeline execution. Traces give you deep visibility into what happens during each call -- every LLM request, every tool invocation, and the latency of each pipeline stage.


Why Tracing Matters

Voice AI pipelines involve multiple components running in real time: speech-to-text, LLM inference, tool execution, and text-to-speech. When something goes wrong -- slow responses, incorrect answers, failed tool calls -- you need visibility into exactly what happened and where the bottleneck occurred.

Tracing provides:

  • LLM call inspection -- see the exact prompts, completions, token usage, and latency for every LLM call.
  • Tool debugging -- view tool invocation inputs, outputs, and duration.
  • Latency analysis -- identify which pipeline stage is causing delays.
  • Cost tracking -- understand per-call LLM and provider costs.
  • Quality monitoring -- review conversation patterns across many calls.

How It Works

When tracing is enabled, zoxaAI exports OpenTelemetry spans to Langfuse for every pipeline run. Each span captures a discrete operation within the call pipeline. Spans are nested into a trace tree that represents the complete execution timeline for a single call.

Trace: Pipeline Run (call_a1b2c3d4)
  |-- Span: LLM Generation (GPT-4.1, 350ms)
  |-- Span: Tool Call (check_appointment, 120ms)
  |-- Span: LLM Generation (GPT-4.1, 280ms)
  |-- Span: TTS Synthesis (ElevenLabs Flash v2.5, 45ms)
  |-- Span: STT Transcription (Soniox stt-rt-v5, 30ms)

Setting Up Tracing

Get Your Langfuse Credentials

If you use Langfuse Cloud:

  1. Go to cloud.langfuse.com.
  2. Select your project (or create one).
  3. Navigate to Settings > API Keys.
  4. Copy the Public Key and Secret Key.

If you self-host Langfuse, use your instance URL as the Host.

Open Organization Settings

In the zoxaAI dashboard, navigate to Settings in the sidebar.

Configure Langfuse Credentials

Find the Langfuse section and enter your credentials:

FieldDescriptionExample
HostYour Langfuse instance URLhttps://cloud.langfuse.com
Public KeyYour Langfuse project's public keypk-lf-...
Secret KeyYour Langfuse project's secret keysk-lf-...

Save

Click Save. All calls from your organization will now be traced to your Langfuse project.

Tracing starts immediately

Once credentials are saved, tracing begins on the next call. Existing completed calls are not retroactively traced.


What Gets Traced

Each pipeline run generates spans for every major component:

ComponentSpan TypeDetails Captured
LLMGenerationModel name, prompt/completion text, tokens in/out, latency, cost
Tools / FunctionsSpanTool name, input arguments, output result, duration
TTSSpanProvider, voice name, characters synthesized, latency
STTSpanProvider, model, transcription result, latency
PipelineTrace (root)Overall call duration, frame count, error status

Every span is tagged with the organization ID (zoxa.org_id) for filtering in Langfuse.


Viewing Traces

After tracing is configured, each call's detail page in zoxaAI includes a Trace URL that links directly to the trace in your Langfuse dashboard.

In Langfuse

Once in Langfuse, you can:

ActionDescription
View span treeSee the complete nested execution timeline for a call
Inspect LLM callsRead the exact prompts, system messages, and completions
Analyze token usageSee prompt tokens, completion tokens, and estimated cost per LLM call
Measure latencyIdentify slow spans (e.g., a tool call that took 2 seconds)
Filter by metadataFilter traces by organization ID, model, or custom tags
Compare tracesCompare execution patterns across different calls
Track costsAggregate LLM costs across all calls

Multi-Organization Routing

zoxaAI supports per-organization Langfuse credentials. Each organization's spans are routed to their own Langfuse project independently.

ScenarioBehavior
Organization A has Langfuse credentials configuredOrganization A's call traces go to Organization A's Langfuse project
Organization B has Langfuse credentials configuredOrganization B's call traces go to Organization B's Langfuse project
Organization C has no Langfuse credentialsOrganization C's traces fall back to the platform-level default exporter (if one exists)

There is no cross-organization data leakage -- each organization's traces are completely isolated.


Platform-Level Configuration

Platform administrators can configure a default Langfuse exporter via environment variables. These serve as the fallback when an organization does not have its own credentials configured.

Environment VariableDescriptionDefault
ENABLE_TRACINGSet to true to enable the tracing subsystemfalse
LANGFUSE_HOSTDefault Langfuse host URL--
LANGFUSE_PUBLIC_KEYDefault Langfuse public key--
LANGFUSE_SECRET_KEYDefault Langfuse secret key--

Platform-level vs. organization-level

Organization-level credentials always take priority over platform-level environment variables. The environment variables are only used as a fallback for organizations that have not configured their own Langfuse project.


Use Cases

Debugging Slow Responses

If users report that the agent takes too long to respond:

  1. Open the call in Call History and click the Trace URL.
  2. Look at the span durations for LLM, TTS, and STT.
  3. Identify which component has the highest latency.
  4. If the LLM is slow, consider switching to a faster model.
  5. If TTS is slow, try a different TTS provider or voice.

Investigating Wrong Answers

If the agent gives incorrect answers:

  1. Open the trace for the problematic call.
  2. Inspect the LLM Generation span to see the full prompt and completion.
  3. Check if the system prompt was correct and if the right context was provided.
  4. Review tool call spans to see if tools returned the expected data.

Cost Optimization

  1. In Langfuse, filter traces by date range and aggregate token usage.
  2. Identify which calls consume the most tokens.
  3. Consider optimizing system prompts to reduce token count.
  4. Evaluate whether a smaller/cheaper model can handle your use case.

Troubleshooting

IssueCauseSolution
No traces appearing in LangfuseCredentials not saved or incorrectVerify the Host, Public Key, and Secret Key in Settings
Traces appear but are incompleteCall ended abnormally (error/crash)Check the call's ended reason in Call History
"Trace URL" not shown on call detailTracing not configured for the organizationAdd Langfuse credentials in Settings
Traces going to wrong projectCredentials point to a different Langfuse projectUpdate the Public Key and Secret Key to match the correct project

On this page