OpenTelemetry Metric Cardinality Limits Explained

Updated on
9 min read

OpenTelemetry metric cardinality limits help keep metric series bounded as an application gains users, routes, and services. Developers and site reliability engineers need to understand both the danger of unbounded attributes and the behavior of an SDK when its limit is reached. This guide explains what cardinality measures, how the OpenTelemetry Metrics SDK applies a limit, what overflow means for dashboards, and how to select dimensions that preserve useful operational detail without turning every request into a new time series.

What Are OpenTelemetry Metric Cardinality Limits?

OpenTelemetry is a set of APIs, SDKs, instrumentation libraries, semantic conventions, and protocols for generating and moving telemetry. The OpenTelemetry project does not provide a metrics database or dashboard; it standardizes how applications describe and export measurements.

A metric instrument, such as a counter or histogram, records measurements with attributes. A combination of attribute values identifies a distinct set of measurements for that instrument. If an HTTP request counter has http.route, http.request.method, and http.response.status_code attributes, a request to /orders/{orderId} using GET with a 200 response contributes to one combination. A request to another route or with a different status contributes to another.

The number of distinct combinations is the metric’s cardinality. It is not the number of samples collected: a single time series can receive many samples over time. The OpenTelemetry metrics data model describes how metric points and their attributes represent those measurements.

A cardinality limit bounds how many distinct data points an SDK collects for an instrument during a collection cycle. It is a guardrail against unbounded attribute sets, not a rule that makes any particular attribute safe or useful.

The Problem Metric Cardinality Limits Solve

Attributes make metrics useful to query. A team can compare request duration by route, method, or response class instead of seeing only one application-wide number. The risk appears when an attribute can take a new value for nearly every request. Examples include a user ID, session token, trace ID, full URL containing identifiers, or arbitrary exception message.

If an instrument has several attributes, the possible combinations multiply. Four methods, six status values, twenty routes, and five regions could produce up to 2,400 combinations, even though only some combinations will occur. Adding a customer or request identifier can make the number grow with traffic rather than with the system’s stable behavior.

Each active series requires work in the application SDK, exporter, Collector, or backend. A large number can consume memory, increase payload and storage volume, expand indexes, slow queries, and raise costs. High-cardinality labels can also make dashboards and alerts harder to reason about because the useful signal is buried in an enormous set of series. Prometheus instrumentation guidance similarly warns against labels with unbounded values.

Reducing the scrape interval does not fix cardinality: it changes how often existing series are sampled, not how many unique series exist. Retention limits can reduce historical storage but do not prevent an application from creating a large number of active combinations.

How OpenTelemetry Metric Cardinality Limits Work

The OpenTelemetry Metrics SDK specification defines a cardinality limit as a hard cap on the number of metric points collected for a single instrument during one collection cycle. The specification recommends that SDKs support configuration and that the default limit be 2,000 when no more specific value is configured. Support and configuration details can vary by SDK, so check the implementation used by a service.

Attribute filtering can happen before the SDK evaluates the limit. This matters because removing an unnecessary dimension first can merge measurements into fewer combinations while preserving the remaining dimensions. A View can select which instruments it applies to and which attributes remain. A limit is still useful as a last line of defense if an overlooked or newly introduced attribute expands cardinality.

When an SDK reaches the limit, OpenTelemetry defines an overflow attribute set containing otel.metric.overflow=true. Measurements that cannot be kept under their original attribute combinations are aggregated into the overflow point rather than silently disappearing from SDK aggregation. That bounds the number of distinct combinations, but it loses the ability to split overflow measurements by their original attributes. A dashboard can still show aggregate activity, but it cannot recover which user, route, or other dimension contributed to that point.

Cardinality controls apply at different stages and protect different resources:

Control Where it acts What it protects Trade-off
Bounded instrumentation attributes In the application SDK, export pipeline, and backend Some request-level detail is intentionally not recorded
Attribute filtering or a View In the SDK, before aggregation SDK and downstream series Removed dimensions cannot be queried later
SDK cardinality limit Per instrument and collection cycle SDK aggregation memory and emitted points Overflow combines otherwise distinct dimensions
Collector or backend filtering After SDK aggregation or export Downstream storage and query load Does not prevent upstream SDK work or traffic
Retention or downsampling In the storage and query layer Historical data volume Does not constrain active series cardinality

The most reliable design combines these controls. Start with bounded dimensions in instrumentation, filter unnecessary attributes before aggregation, and configure a limit appropriate to the service. Downstream limits can provide additional protection, but should not be mistaken for controls on SDK memory or export volume.

Components and Key Concepts

Instruments and attributes: Counters, gauges, and histograms record values. Attributes explain those values, but every added dimension can expand the number of combinations. Prefer stable categories such as a route template over an actual URL path.

Resource attributes: Resource attributes identify the entity producing telemetry, such as service.name or deployment.environment. Some attributes, such as a unique service instance identifier, can also increase the number of backend series. Avoid placing request-specific identities on metric resources or instrument measurements.

Views and SDK limits: Views can select instruments and filter attributes before aggregation. A cardinality limit caps the distinct data points an SDK collects for a selected instrument. These are related but separate controls: filtering reduces combinations by removing dimensions, while a limit handles combinations that remain over the configured bound.

Collectors and backends: A Collector can route or transform telemetry, and a backend can enforce its own quotas. Those components may reduce what is retained, but they receive data after the SDK has produced it. Their controls complement rather than replace appropriate instrumentation and SDK configuration.

Real-World Use Cases

For HTTP services, route templates such as /orders/{orderId} keep route dimensions stable. Recording /orders/84721 as a separate route for every order creates unnecessary series. Method and response status are usually bounded values; query strings, raw paths, and user identifiers are not.

In multi-tenant systems, an organization may need to compare broad service behavior by plan, region, or a small tenant cohort. If operators need a particular customer’s history, logs or traces with appropriate access controls may be a better place for that detail than a metric label on every measurement.

LLM services often measure request duration, token usage, errors, and model behavior. Model names and operation categories can be useful bounded dimensions, while prompts, completions, request IDs, and user IDs can create both privacy risks and high cardinality. The existing OpenTelemetry guide for LLM production observability discusses choosing useful signals without treating every detail as a metric dimension.

Getting Started: Set and Check a Limit

For Java applications, add the OpenTelemetry SDK dependency to a Maven project:

<dependency>
  <groupId>io.opentelemetry</groupId>
  <artifactId>opentelemetry-sdk</artifactId>
  <version>1.44.0</version>
</dependency>

The Java SDK exposes a per-View cardinality setting. This example applies a limit to the http.server.request.duration instrument:

import io.opentelemetry.sdk.OpenTelemetrySdk;
import io.opentelemetry.sdk.metrics.InstrumentSelector;
import io.opentelemetry.sdk.metrics.SdkMeterProvider;
import io.opentelemetry.sdk.metrics.View;

SdkMeterProvider meterProvider = SdkMeterProvider.builder()
    .registerView(
        InstrumentSelector.builder()
            .setName("http.server.request.duration")
            .build(),
        View.builder()
            .setCardinalityLimit(500)
            .build())
    .build();

OpenTelemetrySdk openTelemetry = OpenTelemetrySdk.builder()
    .setMeterProvider(meterProvider)
    .buildAndRegisterGlobal();

Register the provider once in a service that owns SDK setup. Add the application’s existing metric reader or exporter to the provider builder before building if measurements must be exported, and do not replace a provider already installed by an auto-instrumentation agent. This example focuses on the limit rather than backend-specific exporter configuration. Shut the provider down as part of the application’s lifecycle so pending telemetry can be flushed. The example sets an explicit cap for one instrument; it does not guarantee a particular number of stored series across an entire backend, where resource attributes and other instruments also contribute.

Before raising a limit, inspect the attributes on the instrument and remove values that identify individual requests or users. Prefer framework-provided route templates over raw URL paths. A limit that is too low can combine useful dimensions into overflow, while a very high limit may provide little protection against accidental cardinality growth.

After exercising representative traffic, compare the series count in the backend and look for the overflow attribute. For a Prometheus-compatible backend, a query such as the following counts matching active series; the metric name and labels may need adjustment for the exporter and backend:

count({job="checkout-api", __name__=~"http_server_request_duration.*"})

Also review instrument attributes and SDK diagnostics when counts increase unexpectedly. If overflow appears, determine which dimensions are generating combinations before changing the cap. When high-detail per-request investigation is needed, use sampled traces or access-controlled logs rather than adding unique identifiers to every metric.

Common Misconceptions

  • “Cardinality is the same as metric volume.” Cardinality is the number of distinct attribute combinations. Sample volume is how often those combinations receive measurements.
  • “The default limit is a universal backend quota.” The specification describes SDK behavior per instrument and collection cycle. A storage service may have separate project, tenant, or ingestion limits.
  • “Overflow means every measurement is dropped.” The SDK’s overflow point preserves aggregation for measurements that could not be represented under their original attributes. It loses their separate dimension breakdown, and downstream systems may still apply their own dropping or quota rules.
  • “Filtering in the backend fixes the source.” A backend can reduce retained data, but it cannot undo SDK aggregation and export work already performed. Remove unnecessary dimensions as early as practical.

Changelog

  • Initial publication.
TBO Editorial

About the Author

TBO Editorial writes about the latest updates about products and services related to Technology, Business, Finance & Lifestyle. Do get in touch if you want to share any useful article with our community.