OpenTelemetry for LLM Production Observability
A production language-model request can pass through an API, a retrieval system, a reranker, one or more model calls, and tools before returning an answer. A slow or incorrect response may come from any part of that path, while a basic request-duration metric only shows that something took too long. OpenTelemetry for LLM production observability gives teams a vendor-neutral way to connect those operations with traces and measure their behavior without making a particular AI platform the source of truth. This explainer covers the telemetry model, useful instrumentation, and the privacy and cost controls needed to operate it responsibly.
What Is OpenTelemetry for LLM Observability?
OpenTelemetry is an open-source set of APIs, SDKs, instrumentation libraries, semantic conventions, and protocols for creating and exporting telemetry. In an LLM application, it can describe the overall request and the work inside it: retrieving documents, calling a model, invoking a tool, or retrying after a failure. The application can send that data through the OpenTelemetry Protocol (OTLP) to a Collector or compatible observability backend.
OpenTelemetry is not an LLM, an evaluation framework, or a place where telemetry is stored and searched. It standardizes how telemetry is produced and transported; teams still choose a backend and decide what they need to measure. The same application instrumentation can therefore be routed to different systems without tying every model call to one vendor’s monitoring API.
Why LLM Systems Need Production Observability
An ordinary web request often looks like one operation from the service’s point of view. An LLM feature can contain several dependent operations with different failure modes: authorization, prompt construction, cache lookup, embedding generation, retrieval, reranking, model inference, and tool execution. A request can be slow because it waited in a queue, fetched too many documents, retried a provider call, or generated a long response. A single application latency number cannot distinguish those causes.
Model-serving measures also have different meanings. Time to first token helps describe the wait before a streamed answer begins; inter-token latency describes generation pace; total duration includes the full request path. Token counts can support usage analysis, but they do not by themselves represent a bill because provider pricing, cached tokens, model choices, and contractual rates differ. To reason about these measures together, teams need a request-level view as well as aggregated service metrics.
Observability is also useful when an application combines retrieval or tools with generation. If a response cites irrelevant material, a trace can show which passages the retriever returned and whether the model call followed. It cannot prove that the final answer is correct. Evaluation datasets, human review, and application-specific quality checks remain necessary.
How LLM Tracing Works
A trace represents one request, with spans representing timed operations within it. For example, an API request can be the parent span, with child spans for retrieval, a model call, and a tool call. Each span has a name, start and end times, status, and optional attributes or events. The resulting hierarchy helps an operator see where time was spent and which operation failed.
When a request crosses an instrumented service boundary, trace context can be propagated so the receiving service continues the same trace. The W3C Trace Context specification defines HTTP headers such as traceparent and tracestate for this purpose. A model provider may not continue an application’s trace across its own infrastructure, so the client-side span still matters: it records the time and outcome visible to the calling service without implying visibility into the provider’s internal execution.
Traces work alongside other signals, but answer different operational questions:
| Signal or method | What it helps answer | Example for an LLM application | Important limit |
|---|---|---|---|
| Metrics | How often, how much, or how slowly? | Request rate, error rate, duration, or token-use distribution | Aggregation does not explain one particular request |
| Traces | Which steps made this request slow or unsuccessful? | Time in retrieval, retries, model calls, and tool execution | Sampling and missing instrumentation can leave gaps |
| Logs | What event or error was recorded? | Provider timeout or a retrieval service exception | High volume and sensitive content require controls |
| Evaluation | Did the output meet a task-specific quality bar? | Groundedness or expected-answer checks on a test set | Not supplied by telemetry collection itself |
For streaming responses, a span around the full model request records total duration. Separate measurements or carefully chosen span events can capture time to first token and other relevant milestones. Do not turn every generated token into a span; that creates excessive data and does not improve most investigations.
Components and Key Concepts
Instrumentation creates telemetry around the work. Automatic instrumentation can capture supported web frameworks and network clients, while manual spans can describe application-specific steps such as a retrieval stage or agent handoff. Add manual instrumentation where it fills a real gap; avoid creating duplicate spans around operations an existing library already captures.
The SDK supplies runtime behavior such as resources, sampling, span processors, and exporters. A resource identifies the producing service with attributes such as its name and deployment environment. Stable service identity makes it possible to compare traces across instances without attaching request-specific information to every span.
Semantic conventions give commonly used attributes consistent names and meanings. The OpenTelemetry project’s GenAI span conventions describe attributes for model operations, including the operation, provider, and requested model. The project marks these GenAI conventions as development, so check their current status and instrumentation support before relying on a name or treating it as a stable contract.
An optional OpenTelemetry Collector receives telemetry, processes it, and exports it to one or more destinations. A Collector can centralize routing, batching, filtering, and credentials. A backend stores and queries the exported data. Neither the Collector nor OpenTelemetry automatically sets useful alerts, decides which prompts are safe to retain, or explains model quality.
Sampling and data policy determine what leaves the application. Sampling can reduce trace volume, but sampled traces are examples rather than a complete request count. Attributes such as model name and operation are often useful; raw prompts, completions, user identifiers, and tool arguments can expose sensitive data. Define and test redaction before exporting telemetry.
Real-World Use Cases
- Diagnosing latency: Follow a slow request through queueing, retrieval, model inference, and response streaming to identify which stage changed.
- Investigating errors: Connect a failed model span or tool call to its parent request, status, retry, and relevant error event without putting credentials or full payloads in a log.
- Understanding RAG behavior: Compare the retrieved and reranked steps with the subsequent model call. Tracing shows what the application passed along; it does not establish whether those documents support the answer. For retrieval architecture, see hybrid search for RAG.
- Tracking serving changes: Compare latency and token-use distributions by service version or model, using metrics for aggregate trends and traces to inspect representative requests. Changes should be evaluated against the same workload.
- Reviewing tool-using workflows: See which tools were called, in what order, and how much time each took. Record tool names and outcomes, not sensitive arguments or returned content by default.
These techniques complement general OpenTelemetry tracing and signals. For the serving-side performance measures that affect the same requests, see LLM inference optimization and KV caching.
Getting Started: Instrument a Model Operation
The following local setup starts a Collector that accepts OTLP over HTTP and prints received spans. Create otel-collector.yaml:
receivers:
otlp:
protocols:
http:
endpoint: 0.0.0.0:4318
processors:
batch: {}
exporters:
debug:
verbosity: basic
service:
pipelines:
traces:
receivers: [otlp]
processors: [batch]
exporters: [debug]
Create a compose.yaml beside it:
services:
otel-collector:
image: otel/opentelemetry-collector-contrib:latest
command: ["--config=/etc/otelcol-contrib/config.yaml"]
volumes:
- ./otel-collector.yaml:/etc/otelcol-contrib/config.yaml:ro
ports:
- "4318:4318"
Start the Collector, then install the Python SDK and OTLP/HTTP exporter. The Collector configuration documentation explains how receivers, processors, exporters, and signal pipelines fit together; the Python documentation covers SDK and instrumentation choices.
docker compose up -d otel-collector
python -m pip install opentelemetry-api opentelemetry-sdk opentelemetry-exporter-otlp-proto-http
Save this smoke test as trace_demo.py. It exports one model-operation span to the local Collector. The callable is a demonstration stub; in an application, wrap the actual provider SDK call and record only attributes approved by your data policy.
from opentelemetry import trace
from opentelemetry.exporter.otlp.proto.http.trace_exporter import OTLPSpanExporter
from opentelemetry.sdk.resources import Resource
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor
provider = TracerProvider(
resource=Resource.create({"service.name": "llm-api"})
)
provider.add_span_processor(BatchSpanProcessor(OTLPSpanExporter()))
trace.set_tracer_provider(provider)
tracer = trace.get_tracer(__name__)
def traced_model_call(invoke):
with tracer.start_as_current_span("chat example-model") as span:
span.set_attribute("gen_ai.operation.name", "chat")
span.set_attribute("gen_ai.provider.name", "example")
span.set_attribute("gen_ai.request.model", "example-model")
return invoke()
if __name__ == "__main__":
print(traced_model_call(lambda: "demo response"))
provider.force_flush()
provider.shutdown()
Run the test and inspect the Collector output:
python trace_demo.py
docker compose logs --tail=50 otel-collector
The debug output should include the service resource and model-operation span. In a real service, use maintained instrumentation for supported libraries where available, and add manual spans for uncovered work such as retrieval or custom tools. Measure request rate, errors, duration, time to first token, and token usage with metrics; use traces to investigate individual paths. Keep metric labels bounded and avoid user, prompt, or request identifiers as metric dimensions. For model-cost analysis, combine telemetry with actual provider billing data rather than assuming token counts equal spend.
Before sending data to a production backend, verify exporter endpoints, Collector authentication, network access, sampling, retention, and attribute redaction. The local debug exporter is for a smoke test, not a production storage or security configuration. Pin the Collector image to a tested release outside a local experiment, and monitor the Collector pipeline itself so dropped or queued telemetry is visible.
Common Misconceptions
- “OpenTelemetry is an LLM monitoring platform.” It provides APIs, instrumentation, conventions, and transport. A backend is still required to store and analyze telemetry.
- “A model span explains how the provider produced an answer.” A client span describes the operation visible to your application. It does not expose the provider’s internal queue, hardware, or model reasoning.
- “Tracing tells us whether an answer is correct.” It can show inputs and steps only to the extent safely instrumented. Quality needs evaluation methods appropriate to the task.
- “More detail is always better.” Prompts, completions, tool arguments, and user identifiers can be sensitive, expensive, or high-cardinality. Collect only what supports an operational question, and apply access and retention controls.
Related Articles
- OpenTelemetry Signals and Distributed Tracing Explained
- LLM Inference Optimization and KV Caching Explained
- Hybrid Search for RAG: Combining Vector and Sparse Retrieval
- Observability vs Monitoring: A Beginner’s Guide
Changelog and Last Updated
Last updated: September 29, 2026. Initial publication.

