OpenTelemetry Continuous Profiling Explained

Updated on
10 min read

When a service’s CPU use or latency rises, metrics can show the impact and traces can identify slow requests, but neither necessarily reveals which functions consumed the resources. OpenTelemetry continuous profiling aims to connect sampled, code-level performance data with the rest of a service’s telemetry. This explainer covers how profiles work, where the OpenTelemetry Profiles signal fits, and how to approach production profiling without assuming every language or backend supports the same features.

What Is OpenTelemetry Continuous Profiling?

Profiling measures how a running program uses resources by recording samples of its execution, commonly as call stacks with associated values. A CPU profile can show which functions account for sampled processor time. Allocation and heap profiles show different aspects of memory: allocations reveal where memory is created, while a heap profile can show what remains live at a particular point. Wall-clock and contention profiles can help investigate waiting rather than active CPU work.

Continuous profiling collects profiles repeatedly over time, rather than only during a manually triggered diagnostic session. Sampling makes it practical to observe production workloads without recording every instruction or event. It does not make profiling free: sample frequency, stack depth, symbol resolution, runtime behavior, and export volume all affect overhead and data size.

OpenTelemetry provides a vendor-neutral profile data model and ways to carry profiles alongside other telemetry. The OpenTelemetry project maintains the APIs, specifications, and ecosystem. Its Profiles documentation describes profiles as a signal that complements traces, metrics, and logs. The OpenTelemetry Profiles specification is currently marked Alpha, so support and compatibility must be checked for the specific profiler, SDK or agent, Collector distribution, and backend in use.

OpenTelemetry is not itself a profiler or a profiling database. A language runtime, operating-system tool, or profiling agent still has to collect samples. A compatible exporter and backend are needed to transport, store, query, and visualize them.

The Problem Continuous Profiling Solves

Aggregated metrics are good at detecting that a service is using too much CPU, allocating memory quickly, or responding slowly. They generally do not identify the code responsible. A trace can reveal that a particular request spent time in one service or dependency, but a span usually does not contain a sampled call stack for each moment of execution. Logs explain selected events, not the distribution of work across functions.

Without profiles, a team may have to reproduce a problem locally, attach a debugger, or collect a short-lived profile after an alert. Intermittent issues can disappear before an engineer attaches a profiler, and a test environment may not have the same traffic mix as production. Continuous collection provides a historical view of sampled execution that can help narrow an investigation to a function, code path, or resource pattern.

The OpenTelemetry profile model addresses a second problem: profiling tools have often used different formats and identifiers. Standardized data can make it easier to route profile telemetry through a common pipeline and correlate it with service resources or traces. That goal does not mean every profiler emits OpenTelemetry data yet, or that all backends expose identical queries and visualizations.

How Continuous Profiling Works

A profiler periodically samples a running thread or runtime and records the stack at that instant. The sample is associated with a value, such as CPU time, bytes allocated, or a count. A collection of weighted stacks can then be aggregated into a view such as a flame graph. The width of a frame in a flame graph represents its share of the selected sample value, not necessarily its exact wall-clock duration.

A typical telemetry path has several stages:

  1. A runtime, host profiler, or agent samples execution and gathers stack information.
  2. The collector resolves or records symbol and mapping information so machine addresses can be connected to functions and binaries.
  3. The integration adds service and process resource attributes, such as service name and deployment environment, and maps data to a supported profile format.
  4. An exporter sends profiles to a compatible endpoint, directly or through a Collector where that signal and pipeline are supported.
  5. A profiling backend stores samples and provides queries over time, code, service, and selected attributes.

The OpenTelemetry Profiles data model can represent stack samples and mappings, and a sample may include a link to a trace and span. That link can let an engineer move from an expensive code path to the request context in which it was observed. It is optional and depends on the instrumentation and profiler being able to associate the two; a profile is not automatically a complete trace of every request.

Profiles answer a different question from the other common telemetry signals:

Signal Typical data Best question Main limitation
Metrics Aggregated numeric measurements over time Is resource use or service behavior changing? Aggregation usually does not identify the responsible code path
Traces Timed spans for individual operations Where did this request or operation spend time? Sampling and instrumentation gaps can hide work
Logs Timestamped event records What event or error was recorded? Events are not a statistical view of all execution
Profiles Sampled stacks with values such as CPU time or bytes Which functions account for sampled resource use? Interpretation depends on profile type, symbols, and sampling coverage

The OpenTelemetry specification describes how profile samples can reference trace and span context. Correlation is valuable, but sampling policies may mean a profile and trace do not cover the same population. Use metrics for complete aggregate rates when appropriate, traces for request paths, logs for events, and profiles for code-level resource attribution.

Components and Key Concepts

Profile type: The type determines what a sample means. On-CPU profiles sample code while a thread is running and are useful for CPU hotspots. Wall-clock profiles can include time when a thread is waiting, depending on the profiler. Allocation profiles record memory creation activity; heap profiles describe live memory at a point or over an interval. Lock or contention profiles focus on synchronization waits. A runtime may support only some of these types.

Sampling and overhead: Sampling frequency controls how often stacks are captured. A higher frequency can provide more detail but may increase CPU cost, profile volume, and backend storage. A low frequency can miss short-lived work. Measure overhead under representative traffic and choose a rate that answers the operational question rather than maximizing data collection.

Stack traces and symbols: A sample’s value is useful only if its stack can be interpreted. Native binaries may need debug symbols, and compiled or JIT runtimes may require runtime-specific symbol support. Builds, container images, and symbol files should be managed so profiles can be resolved to the correct code version.

Resources and attributes: Stable service identity and deployment attributes let teams compare profiles by service, version, environment, or host. Avoid high-cardinality or sensitive values such as request bodies, credentials, or user identifiers. Profile data can reveal function names, execution patterns, and workload behavior, so it needs access and retention controls.

Export and storage: The profiler, data format, exporter, Collector, and backend must agree on supported profile types and transport. OpenTelemetry standardizes parts of the data model, but it does not guarantee that every Collector build has profile receivers and exporters or that every backend supports the full model. Confirm the complete path before rolling it out.

Real-World Use Cases

  • CPU hotspots: Compare profiles before and after a change to find functions consuming a large share of sampled CPU time.
  • Memory growth: Use allocation profiles to locate high-volume object creation, or heap profiles to inspect objects retained in memory. These answer different questions and should not be treated as interchangeable.
  • Latency investigations: Start with a trace or latency metric to find an affected service, then use a time-aligned profile to investigate expensive code paths. A span-to-profile link can shorten this transition when both signals are supported.
  • Lock contention: Profile synchronization waits to find code that serializes work, especially when request latency rises while CPU use remains comparatively low.
  • Production regressions: Compare profiles across releases, environments, or workload periods to identify changes in execution costs without relying only on a reproduction in a development environment.

Profiling is most useful when it is combined with a clear question. A CPU profile cannot explain a database outage by itself, and a heap profile may not identify why a request is slow. Pair profiles with the signal that exposed the symptom and verify that the profile window, service version, and workload actually match the incident.

Getting Started with Continuous Profiling

Begin with one service and one profile type. Check whether the runtime, host, or agent supports that profile and whether it can export a format accepted by the chosen backend. Because the OpenTelemetry Profiles specification is Alpha, verify the versions and limitations in the official Profiles documentation before designing an end-to-end OpenTelemetry pipeline.

On a Linux test host, perf provides a simple way to collect and inspect a short CPU profile. Install the perf package supplied for the host’s distribution and kernel, then check that it is available:

perf --version

Record a 30-second sample at approximately 99 samples per second, with call graphs enabled:

sudo perf record -F 99 -g -- sleep 30

Inspect the resulting profile:

sudo perf report

This example creates a local perf.data file for analysis. It demonstrates sampled profiling; it does not export data as OpenTelemetry or configure continuous collection. For an OTel deployment, choose an integration that explicitly supports the profile signal, confirm how it resolves symbols and assigns service resources, and verify every hop through the exporter, Collector, and backend.

Roll out gradually. First compare service-level CPU and latency before and during profiling. Check that reports contain readable stacks and the expected service version; unknown frames often point to missing or mismatched symbols. Then validate trace/profile correlation if the integration claims to provide it. Set retention and access policies, filter sensitive attributes, and monitor dropped samples, exporter failures, backend ingestion, and the profiler’s own resource use. If a host denies access to perf, check distribution permissions and kernel policy rather than weakening production security controls.

Common Misconceptions

“Continuous profiling records every line of every request.” Profilers usually sample execution. A profile is statistical evidence over a time window, not a deterministic record of each instruction or request.

“A profile is just another trace.” A trace describes timed operations and their relationships. A profile aggregates stack samples and values across execution. A profile sample can be linked to trace context, but neither signal replaces the other.

“OpenTelemetry automatically profiles every instrumented service.” OpenTelemetry provides a model and ecosystem for telemetry; it does not turn on a profiler by itself. Runtime support, compatible integrations, Collector components, and backend features all need to be present.

“A profile with a trace link explains the whole request.” The link can add useful context, but both signals may be sampled or incomplete. A profile can point to a costly stack without containing every operation in the trace, and a trace may not have a corresponding profile.

Changelog and Last Updated

Last updated: October 4. Initial publication.

TBO Editorial

About the Author

TBO Editorial writes about the latest updates about products and services related to Technology, Business, Finance & Lifestyle. Do get in touch if you want to share any useful article with our community.