OpenTelemetry Signals and Distributed Tracing Explained

Updated on
10 min read

When a request crosses an API gateway, several services, and a database, a dashboard can show that it is slow without showing where the delay occurred. OpenTelemetry gives application and platform teams a vendor-neutral way to create and move telemetry across those boundaries. This explainer covers its three core signals—traces, metrics, and logs—with a closer look at distributed tracing, context propagation, and a practical local Collector setup.

What Is OpenTelemetry?

OpenTelemetry, often shortened to OTel, is a set of APIs, software development kits (SDKs), instrumentation libraries, semantic conventions, and protocols for generating, collecting, and exporting telemetry. It is an open project under the Cloud Native Computing Foundation. The OpenTelemetry project site describes the project, while its official documentation explains how to instrument services and route their data.

OpenTelemetry is not a monitoring dashboard or a storage system. It defines common ways for software to describe and transmit telemetry so an application can send useful data to a compatible backend without binding its instrumentation to one vendor. A backend still has to store, query, visualize, and alert on that data.

The project groups telemetry into three signals:

  • Traces describe the path and timing of an operation as it moves through one or more components.
  • Metrics are measurements aggregated over time, such as request counts, error rates, or queue depth.
  • Logs are timestamped records of events, often containing details about a particular request or failure.

The signals answer different questions. Metrics can reveal that latency has risen across a service; a trace can identify the slow dependency on one request; a correlated log can provide the error message emitted by that dependency.

The Problem OpenTelemetry Solves

Distributed applications are assembled from components that may use different languages, libraries, deployment platforms, and observability products. Without shared conventions, every team may choose different agent APIs, trace formats, attribute names, or export mechanisms. Operators then have to maintain custom adapters and often lose context as a request crosses service boundaries.

OpenTelemetry provides common instrumentation interfaces and the OpenTelemetry Protocol (OTLP), so services can emit telemetry through a consistent pipeline. Instrumentation can be added in application code, supplied by libraries, or provided by an automatic agent. Teams can change exporters or route data through a Collector without rewriting every business operation around a backend’s proprietary SDK.

This standardization does not make telemetry automatically complete or correct. Teams still decide which operations matter, which attributes are safe, how much data to retain, and which alerts or queries are useful. OTel reduces integration friction; it does not replace operational design.

How OpenTelemetry and Distributed Tracing Work

A trace represents one logical operation, such as completing a checkout. It is made of spans: timed records for units of work such as handling an HTTP request, querying a database, or calling another service. Spans have identifiers, start and end times, attributes, events, and status. A child span refers to its parent so a tracing backend can reconstruct the request tree.

For a trace to continue across processes, the caller has to propagate trace context with the outgoing request and the receiver has to extract it. The W3C Trace Context specification standardizes HTTP headers including traceparent and tracestate. OpenTelemetry propagators can read and write this context. When services use compatible propagation, independently instrumented components can join their spans into the same trace rather than creating disconnected fragments.

An instrumented application typically uses this path:

  1. A library or application instrumentation creates a span around an operation and attaches relevant attributes.
  2. The language SDK applies resource information, sampling, and span processors, then exports completed data.
  3. The exporter sends telemetry, often using OTLP, either directly to a backend or to an OpenTelemetry Collector.
  4. A backend stores the data and provides search, visualization, correlation, and alerting.

The Collector is an optional, separately deployable service that receives telemetry, processes it, and exports it to one or more destinations. A pipeline is configured for a signal and connects receivers, processors, and exporters. Teams commonly use a Collector to centralize credentials, batch data, redact attributes, route to multiple backends, or keep backend-specific configuration out of application deployments. It is an important control point, not a guarantee that data cannot be lost: queues, retries, resource limits, and downstream availability still matter.

Metrics, logs, and traces can be compared by the operational questions they answer:

Feature Metrics Logs Traces
Typical shape Aggregated numeric measurements over time Individual timestamped events Timed spans arranged into a request path
Best question How often or how much? What event or error was recorded? Where did this particular operation spend time?
Common strength Efficient dashboards and alert conditions Detailed event context Cross-service latency and dependency relationships
Main limitation Aggregation hides individual requests Volume and inconsistent fields can make investigation difficult Sampling and missing propagation can leave incomplete paths
Useful correlation Service, route, and time window Trace or request identifier Trace and span identifiers

The signals complement one another, but they do not have identical data models or cost profiles. A team can use OpenTelemetry traces and logs while retaining an existing metrics backend, or adopt only the signals that address its current operational gaps.

Components and Key Concepts

Instrumentation is the code that observes operations. Manual instrumentation is useful for business steps that libraries cannot infer, such as a payment authorization or a queue-processing decision. Automatic instrumentation can capture common framework and client-library operations with little application-code change. Both approaches can be used together, but duplicate instrumentation should be avoided.

An API is the interface application code uses to create telemetry. An SDK implements the pipeline behavior for a language, including sampling and export. Instrumentation libraries and agents can use those interfaces so application teams do not need to hand-build span creation for every HTTP or database call.

A resource describes the entity producing telemetry, for example service.name, service.version, and deployment environment. Resource attributes let a backend distinguish checkout-api from inventory-api even when both export the same kind of span. Keep identity attributes stable and avoid placing request-specific values, such as a user ID, on the resource.

Attributes describe a span or event, while semantic conventions provide shared names and meanings for common concepts such as HTTP methods and database operations. Consistent conventions make queries more portable. They do not ensure every library has the same coverage, so verify which attributes your instrumentation actually emits.

Context propagation carries trace identity between work units. HTTP frameworks often propagate it automatically when instrumented, but message queues and custom transports need compatible instrumentation or explicit context handling. Baggage can carry additional key-value context, but it is propagated data—not a trusted authorization signal. Do not place secrets or personal information in baggage.

Sampling limits the volume of recorded traces. Head sampling decides early, before the full request outcome is known. Tail sampling can make a decision after spans arrive, allowing rules based on errors or latency, but requires a Collector setup that keeps spans for the same trace together and has enough memory. Whichever approach is used, document its effect: a sampled trace set cannot be treated as a complete count of all requests.

Real-World Use Cases

  • Finding a slow dependency: A latency metric identifies a service with a rising response time; a trace shows whether the time was spent in the service itself, a database, or a downstream API.
  • Following a request across microservices: Propagated context connects gateway and service spans, making fan-out, retries, and asynchronous work easier to inspect. The broader microservices architecture guide explains why these boundaries make diagnosis more demanding.
  • Correlating incidents: A trace identifier in structured logs can connect a failed span to the application error that caused it. Metrics can then show whether the problem is isolated or widespread.
  • Changing observability backends: A team can keep its application instrumentation stable while changing the Collector’s exporters or routing telemetry to a second destination during a migration.

OpenTelemetry also supports visibility into machine-learning and AI application workflows. Teams can instrument model calls, retrieval steps, and tool invocations, but should carefully filter prompt content and other sensitive values before exporting telemetry.

For a focused walkthrough of those model, retrieval, and tool-call spans, see OpenTelemetry for LLM Production Observability.

Getting Started with OpenTelemetry

The exact SDK installation depends on the application’s language and framework. Start with the official OpenTelemetry instrumentation documentation to select a supported SDK or automatic instrumentation package. The following local example starts a Collector that accepts OTLP and prints received traces through its debug exporter.

Create otel-collector.yaml:

receivers:
  otlp:
    protocols:
      grpc:
        endpoint: 0.0.0.0:4317
      http:
        endpoint: 0.0.0.0:4318

processors:
  batch: {}

exporters:
  debug:
    verbosity: basic

service:
  pipelines:
    traces:
      receivers: [otlp]
      processors: [batch]
      exporters: [debug]

The Collector’s configuration documentation describes this receiver-processor-exporter pipeline. Add a compose.yaml beside the configuration:

services:
  otel-collector:
    image: otel/opentelemetry-collector-contrib:latest
    command: ["--config=/etc/otelcol-contrib/config.yaml"]
    volumes:
      - ./otel-collector.yaml:/etc/otelcol-contrib/config.yaml:ro
    ports:
      - "4317:4317"
      - "4318:4318"

Start the Collector:

docker compose up -d otel-collector

Install and configure an OpenTelemetry SDK or auto-instrumentation agent in the application before setting these standard environment variables. For an app running on the same host, an OTLP/HTTP SDK can use:

export OTEL_SERVICE_NAME=checkout-api
export OTEL_EXPORTER_OTLP_ENDPOINT=http://localhost:4318
export OTEL_EXPORTER_OTLP_PROTOCOL=http/protobuf

These variables configure the service identity and OTLP endpoint for SDKs that support the corresponding environment settings; they do not install instrumentation by themselves. If the application runs in the same Compose network, use http://otel-collector:4318 instead of localhost. Exercise an instrumented endpoint, then inspect received span output with:

docker compose logs -f otel-collector

No spans will appear until the application has an SDK or agent configured to create and export them. If the Collector does not start, check that the mounted file path is correct and that the configuration is valid for the chosen Collector distribution. If it starts but receives nothing, confirm the application’s exporter protocol and endpoint, check host-versus-container networking, and look for SDK export errors. In a production pipeline, pin the Collector image to a release, protect receiver endpoints, size memory and queues for expected traffic, and configure retries and monitoring. Review attributes for secrets and personal data before routing telemetry outside the application boundary.

Common Misconceptions

“OpenTelemetry is a backend.” It creates and moves telemetry; a compatible backend still has to store and query it. The Collector is also not a trace database or user interface.

“Installing an agent produces a complete trace automatically.” Auto-instrumentation covers supported libraries and frameworks, not every business operation or custom transport. Context can be lost at uninstrumented boundaries, and asynchronous workflows need explicit propagation.

“More telemetry always means better observability.” Unbounded attributes, excessive span events, and high-volume logs increase storage, network, and privacy risks. Set useful attributes, limit sensitive data, and choose sampling and retention based on the questions operators need to answer.

“A trace is a complete record of every request.” Sampling may intentionally discard traces, and instrumentation gaps can create partial ones. Use metrics for aggregate rates and alerting; interpret traces as diagnostic examples unless the pipeline is configured to record every operation.

Changelog and Last Updated

Last updated: September 29. Initial publication.

TBO Editorial

About the Author

TBO Editorial writes about the latest updates about products and services related to Technology, Business, Finance & Lifestyle. Do get in touch if you want to share any useful article with our community.