Guide
LLM gateway observability: traces, logs and export
Short answer
An LLM gateway fits an existing observability stack when it exports the signals you need, on an acceptable plan, with usable request context. OpenTelemetry support is a starting point: verify transport, span attributes, sampling, content handling and application trace correlation. “Not documented” is a research gap, not proof that export is impossible.
What OpenTelemetry support establishes
In this catalogue, 20 of 31 products record native, partial or integration-based OpenTelemetry support. Those grades describe an export mechanism, not equivalent spans or a verified end-to-end agent trace. Read the recorded tracing note and source for each candidate.
Two targeted documentation checks on September 16, 2026 corrected earlier catalogue gaps: Vercel Trace Drains require Pro or Enterprise and have separate delivery/egress charges; their payload excludes prompt and completion text. Cloudflare documents OTLP export with prompt/completion attributes. Content policy must therefore be checked independently of the OTel label. These were documentation checks, not collector tests.
For other rows, use the field-check date rather than treating these two corrections as a fresh review of the whole catalogue. Native export, a third-party integration and SDK instrumentation can all be useful, but they place different configuration and maintenance work on your team.
When export is proprietary or not documented
The catalogue records 2 proprietary tracing implementations and 9 rows without a documented tracing answer. Do not assume either group needs a manual polling bridge. Ask about supported exports, webhooks, SDK instrumentation and commercial-plan features; validate the answer with a sample payload.
If only a vendor console is available, decide whether its diagnostics meet your needs. If a bridge is required, budget its latency, correlation gaps and maintenance. Instrumenting your application still helps, but cannot expose unreported internal gateway attempts.
A gateway that tells the caller what it did
The most useful observability feature in this catalogue is not a dashboard. Braintrust Gateway reports failover in the response headers, naming the endpoint that was tried and the one that served the request.
Per-request routing evidence can be attached to an application span while user, tenant and feature context are available. A vendor dashboard may also offer trace-ID correlation or exports; verify that capability rather than assuming timestamp matching is the only option. Ask which provider and model actually served each request, and how to join that record to application telemetry.
This is worth the attention because silent degradation can escape ordinary status-code alerts. Everything returns 200. The answers come from a cheaper model, or an older one, or a different region, and they are slightly worse in a way no alert threshold describes. Latency drifts up and stays inside the budget. Where no header carries that fact, find the log field that does and confirm you can group by it before an incident makes it urgent. What the gateway does when an upstream fails, and how much of that behaviour you can configure, is the subject of the failover guide rather than this one.
What you can see is bounded by what is stored
Tracing tells you the shape of a request. Whether you can reconstruct it depends on a different pair of fields. By default, 11 of 31 products log full content, 12 log metadata only, 6 log nothing until you configure it, and 2 record the question as not applicable to how they are deployed. On what is kept, 14 products make the content stored configurable, 11 keep metadata only, 4 keep the full request and response, and 2 keep nothing.
Logging and trace export can have different content policies. A metadata-only request log does not prove that exported spans exclude prompts, and a trace without prompt text can still expose sensitive custom attributes. Review each destination, collection toggle, sampling rule, access role and retention policy separately. Missing export documentation is not proof that a payload or trace is unavailable.
Choose content collection by environment and sensitivity. Synthetic staging traffic can help debug instrumentation without copying production prompts. If production content is necessary, define authorization, redaction and retention at both the gateway and collector. The prompt-logging guide maps these separate data paths.
Evaluation hooks and prompt management are separate from tracing. 10 of 31 products document evaluation hooks and 12 document prompt management. If your plan for regression testing depends on replaying production traffic, those two fields matter as much as the tracing column, because a trace you cannot replay against is a record rather than a tool.
A trace acceptance test
Illustrative trace tree: application request → gateway request → provider attempt A (timeout) → provider attempt B (success). Tool calls and later model steps belong under the application trace; a gateway request span alone does not reconstruct them.
- Record version, edition, exporter protocol, sampling and content settings.
- Send a synthetic request with trace context; confirm parent/child relationships at the collector.
- Force a safe staging fallback and inspect attempted versus served model, status, duration and usage.
- Repeat with streaming, cancellation and collector failure; inspect both latency and missing spans.
- Check exported attributes for secrets or content and reconcile delivery volume with the billed meter.
Save the payload and observations. This procedure is proposed validation, not evidence that every product passes.
Tracing, logging and streaming, product by product
All 31 products, sorted by name. Read the tracing column as the export path and the two logging columns as separate recorded policies, not the complete contents of exported traces.
A blank reads as not published rather than as no. A vendor who has never written down whether spans leave the process may well export them; the absence records that the documentation does not say so, which is what you can check today. Every row links to the note and the vendor page it was read from.
Streaming is the quiet complication in that last column. 26 of 31 products stream, 4 stream on some endpoints and not others, and 1 publishes nothing either way. Once a response streams, one duration becomes two measurements: time to first token, which is what a user experiences, and total duration, which is what the work cost. A gateway reporting only one of them is hiding the more interesting half, and the latency overheadLatency overhead: The delay the gateway itself adds, separate from the model’s own thinking time. Usually single-digit milliseconds and irrelevant next to a model taking several seconds — treat large claimed differences with suspicion. the gateway itself adds is only visible in the first. Check also that token counts arrive at all on a streamed response, because on several products usage figures come in a final frame that a client can drop without noticing.
Where you sit
OpenTelemetry support is essential if
- You already run a collector and an incident timeline people trust.
- The gateway is one hop in a request that crosses several services.
- Your workload is an agent, and you need spans across the whole run.
- On-call correlates by trace ID rather than by timestamp.
The vendor console is enough if
- Model calls are the only thing you are watching closely.
- Nobody has a collector to send spans to yet.
- You want cost per model on day one without instrumenting anything.
- One team owns the whole path and can watch two tabs.
Check before you wire it in
- Whether spans cover a whole agent run or one call at a time.
- Whether tracing is on by default or behind a plugin flag.
- What the response headers tell the caller per request.
- Whether content storage can be set per environment.
Common questions
Does OpenTelemetry support mean my existing dashboards will just work?
It means the traces can reach your collector, which is the hard part, and not that the spans contain what you want. The catalogue records an OTLP endpoint or exporter as documented support; the note on each product says what the spans actually cover. Several products emit request-scoped spans only, so a single call is visible and a multi-step agent run is not reconstructed for you. Send a real request to a staging collector and read the span attributes before you plan a migration around them.
Is a vendor console instead of OpenTelemetry a problem?
Only if you already run somewhere else to look. A proprietary console is usually better on day one, because it is populated and shaped for model traffic without any work from you. It becomes a problem when an incident spans your application and the gateway, because you cannot join two systems on a trace ID by eye. If the gateway is one hop in a longer request, the export path matters more than the dashboard.
Why does a gateway that logs metadata only make debugging harder?
Latency and token counts tell you that a request was slow or expensive. They do not tell you what was sent, so a wrong answer cannot be traced back to the prompt that produced it. That is the trade-off with the privacy position: the same absence that keeps prompts out of a third party removes the evidence you would want at three in the morning. Decide which risk you are optimising for, per environment rather than globally.
What should I look for in response headers?
Anything that tells the caller what the gateway did on this request: which upstream served it, whether a fallback fired, whether the answer came from a cache. A header can be attached to your own span at the point where you already have context, so it needs no polling and no second system. Where no such header exists, find the log field that carries the same fact and confirm you can group by it before you need to.
How does streaming change the numbers I should be watching?
It splits one measurement into two. Time to first token describes what a user perceives; total duration describes what the model and the network cost you. A gateway that reports only the second hides the half that people complain about, and a gateway that reports only the first hides the half that shows up on the bill. Check which one the metric on the dashboard is, and whether usage figures arrive at all on a streamed response.