OpenTelemetry Collector Pipelines and Tail Sampling Explained
When a production incident affects only a fraction of requests, collecting every trace can be expensive, but sampling requests before their outcome is known can discard the evidence operators need. OpenTelemetry Collector pipelines provide a shared place to receive, process, and export telemetry; tail sampling lets a trace pipeline make retention decisions after spans arrive. This guide explains how the pieces fit together, what the trade-offs are, and how to start with a local configuration.
What Are Collector Pipelines and Tail Sampling?
The OpenTelemetry project provides vendor-neutral APIs, SDKs, instrumentation, and protocols for telemetry. The Collector is a separate service that receives telemetry, applies configured processing, and exports the result to one or more destinations. Its official documentation describes how its components and pipelines are assembled.
A pipeline is a signal-specific path through the Collector. A traces pipeline, for example, connects one or more receivers to processors and exporters. A receiver accepts data using a protocol or from a source; processors can batch, filter, enrich, or sample it; exporters send it to a backend or another Collector. A configured component does nothing unless it is included in a service pipeline.
Tail sampling is a traces processor that waits for spans belonging to a trace, then makes a keep-or-drop decision using rules such as whether the trace contains an error or exceeds a latency threshold. Because it evaluates the assembled trace rather than just its first span, it can preserve unusual, useful traces while discarding routine traffic. The tail sampling processor reference documents supported policies and configuration options.
The Problem These Components Solve
Applications and infrastructure often emit telemetry in different formats and send it to different systems. Putting a Collector between services and their backends centralizes transport and processing: teams can batch data, change exporters, apply redaction, or route signals without putting every backend credential and decision in application code.
Tracing creates a separate cost and diagnostic challenge. A trace might contain spans from an API, several downstream services, and a database. Head sampling in an SDK makes its choice near the start of the request, when it may not yet know that the request will fail or take several seconds. If it drops the trace then, a Collector cannot recover the missing spans later.
Tail sampling moves that choice downstream. The Collector can retain every trace matching a rule while sampling only a fraction of ordinary traces. This is useful when rare failures or latency outliers matter more than a complete record of every successful request. It is not free: the Collector must hold spans in memory while the trace is being evaluated, and the added wait increases the time before retained traces are exported.
How a Pipeline and Tail Sampling Work
A trace pipeline typically follows this sequence:
- An instrumented application sends spans, commonly using OTLP, to a Collector receiver.
- A memory limiter protects the process from accepting more data than it can handle.
- The tail sampling processor groups spans by trace ID and waits for its decision window.
- The processor evaluates configured policies against each trace.
- A batch processor groups retained spans, then an exporter sends them onward.
Processor order matters. Tail sampling needs the spans before they are discarded, and the batch processor should usually follow it so it batches the retained data. Receiver, processor, and exporter names are also distribution-dependent: a processor must be included in the Collector binary or container image being run.
Tail sampling differs from head sampling in where it decides and what information it can use:
| Feature | Head sampling | Tail sampling |
|---|---|---|
| Decision point | In the SDK, when the trace begins | In a Collector, after spans arrive |
| Available evidence | Early context and configured probability or parent decision | Collected spans, including status and observed duration |
| Resource cost | Lower downstream volume and little central buffering | Collector memory for traces awaiting decisions |
| Export timing | Sampled spans can be exported as they finish | Traces are delayed until a decision is made |
| Best fit | Broad reduction where all requests need a consistent early rule | Keeping rare errors or slow traces while reducing routine volume |
Tail sampling does not mean the Collector has a complete view by default. If an SDK has already dropped spans, those spans will never reach the processor. If several Collectors independently run tail sampling, all spans for a trace must be routed to the same sampling instance. In a horizontally scaled gateway tier, use trace-ID-aware routing before the sampling processors; randomly distributing spans can split a trace and produce incorrect decisions. The W3C Trace Context specification standardizes trace context propagation in HTTP headers, but propagation alone does not perform that Collector-side routing.
The decision_wait setting controls how long the processor waits before deciding. A longer window can accommodate slower spans but consumes memory for longer and adds export delay. num_traces sets the number of traces the processor can hold, not a throughput target. Size these values against observed trace arrival rates and bursts, and monitor the Collector for refused data, dropped spans, and memory pressure.
Components and Key Concepts
- Receivers accept telemetry. OTLP receivers can listen for gRPC and HTTP traffic; endpoint binding and network access should be limited to the services that need to export.
- Processors transform or select data. Examples include
memory_limiter,batch, filtering processors, andtail_sampling. Declaring a processor underprocessorsdoes not activate it until a pipeline references it. - Exporters deliver telemetry to a backend or another service. A debug exporter is useful for a local smoke test, not a durable production destination.
- Extensions and connectors add service capabilities or connect pipelines where needed. They are separate from the receiver-processor-exporter sequence and are configured according to their component type.
- Policies define tail-sampling decisions. Most configured keep policies are combined so a trace matching any one of them is retained. Policy choice should reflect a diagnostic goal, not just a desire to minimize volume.
Trace identity is essential. Instrumented services must propagate context across process boundaries so their spans share a trace ID. When a Collector tier is scaled, a trace-aware routing layer must then bring those spans together at one tail sampler. The W3C standard describes trace context transport; it does not guarantee that a message broker, proxy, or Collector load balancer preserves trace affinity.
The distinction between SDK and Collector sampling also matters operationally. SDK head sampling is a useful first reduction point when its rule is appropriate, but it can prevent the Collector from seeing an entire request. Keep upstream SDK sampling broad enough for the tail policies to work, and document the combined effect when both stages sample.
Real-World Use Cases
- Incident retention: Keep traces with error status so investigations can inspect failures that represent a small share of total requests.
- Latency outliers: Retain traces above a duration threshold to help identify expensive database calls, retries, or slow downstream services.
- Cost-aware baseline visibility: Preserve all traces matching high-value policies while retaining a small probabilistic sample of ordinary traffic for routine analysis.
- Centralized routing: Use a Collector to apply shared policies, batch data, and send telemetry to a chosen backend without embedding backend-specific exporters in every service.
Sampling changes the population visible in a trace backend. A dashboard that counts sampled traces cannot be treated as a count of all requests. Use metrics for complete aggregate rates where appropriate, and use sampled traces to investigate representative or deliberately selected examples. Teams already using the OpenTelemetry signals and distributed tracing model can add a Collector sampling tier without changing what a span represents.
Getting Started with Tail Sampling
Use the Collector Contrib distribution for this example because it includes the tail sampling processor. Create otelcol.yaml:
receivers:
otlp:
protocols:
grpc:
endpoint: 0.0.0.0:4317
http:
endpoint: 0.0.0.0:4318
processors:
memory_limiter:
check_interval: 1s
limit_mib: 512
tail_sampling:
decision_wait: 10s
num_traces: 10000
policies:
- name: retain-errors
type: status_code
status_code:
status_codes: [ERROR]
- name: retain-slow-traces
type: latency
latency:
threshold_ms: 1000
- name: keep-baseline-sample
type: probabilistic
probabilistic:
sampling_percentage: 5
batch: {}
exporters:
debug:
verbosity: basic
service:
pipelines:
traces:
receivers: [otlp]
processors: [memory_limiter, tail_sampling, batch]
exporters: [debug]
The rules retain traces with an error, traces over one second, and a five-percent sample of other traces. The policies are alternatives, so the baseline rule does not limit the error and latency rules to five percent. Adjust the policy set to match the data and retention requirements; sampling attributes can contain sensitive information and should be reviewed before export.
Place the config beside this compose.yaml:
services:
collector:
image: otel/opentelemetry-collector-contrib:latest
command: ["--config=/etc/otelcol-contrib/config.yaml"]
volumes:
- ./otelcol.yaml:/etc/otelcol-contrib/config.yaml:ro
ports:
- "127.0.0.1:4317:4317"
- "127.0.0.1:4318:4318"
Validate the configuration and start the service:
docker compose run --rm --no-deps --entrypoint otelcol-contrib collector validate --config=/etc/otelcol-contrib/config.yaml
docker compose up -d
docker compose logs -f collector
Configure an instrumented application to send OTLP traces to http://localhost:4318 using the HTTP/protobuf protocol, or to port 4317 using gRPC. The application needs a compatible OpenTelemetry SDK or agent; starting the Collector alone does not generate spans. Make a test request and look for retained trace output in the Collector logs. A quiet log can mean no instrumentation is exporting, all traces missed the configured policies, or the application cannot reach the receiver.
For production, pin the image to a tested release instead of latest, expose receiver ports only on trusted networks, monitor Collector health and memory, and size num_traces for actual traffic. If scaling the sampling tier, add trace-ID-aware routing before the tail sampler. Also assess retry and queue behavior at exporters, because tail sampling cannot prevent loss when a downstream backend remains unavailable.
Common Misconceptions
“Tail sampling keeps every error trace automatically.” Only spans that reach a tail sampler can be evaluated. Earlier SDK sampling, broken propagation, processor overload, or routing a trace across multiple samplers can leave a trace incomplete or invisible.
“A longer decision window always improves sampling.” It gives late spans more time to arrive, but also increases memory use and delays exports. Choose a window based on observed trace completion and late-span behavior, then test under realistic load.
“Sampling replaces metrics.” Sampling intentionally removes traces. It is useful for cost-aware diagnosis, not for exact request totals. Retain appropriate metrics for aggregate rates and interpret trace data according to the sampling policy.
Related Articles
- OpenTelemetry Signals and Distributed Tracing Explained
- OpenTelemetry Metric Cardinality Limits Explained
- Cloud Observability Tools: Logs, Metrics, Traces, and Choosing a Stack
- Infrastructure Monitoring with Prometheus
Changelog and Last Updated
Last updated: October 2. Initial publication.

