eBPF Observability: How Kernel Tracing Works

Updated on
9 min read

When a Linux service is slow or dropping network traffic, application logs may show the symptom without exposing what the kernel was doing. eBPF observability lets operators run small, event-driven programs at selected points in the Linux kernel and collect focused measurements without rebuilding the application or kernel. This guide explains how that works, what it can and cannot reveal, and how to run a first trace safely.

What Is eBPF Observability?

eBPF is a Linux kernel technology for loading and running restricted programs in response to events. Its name comes from “extended Berkeley Packet Filter,” but modern eBPF is not limited to packet filtering. Programs can observe or act at supported kernel hooks for system calls, networking, schedulers, security decisions, and other operations. The eBPF project’s introduction describes the technology and its safety model.

For observability, an eBPF program usually records selected facts—such as a process name, a duration, a return value, or a packet count—and sends aggregated or event data to a user-space tool. That data can help answer questions about kernel and network behavior that application-level instrumentation cannot see. eBPF does not automatically produce dashboards, distributed traces, or a complete recording of system activity; a separate tool must load programs, collect their output, and present it.

Why eBPF Exists

Traditional diagnostics each expose only part of a system. Application instrumentation can identify a slow function or database call, but it may not explain time spent waiting for CPU, disk, or network activity. Packet captures show traffic visible at a capture point, but not necessarily which process caused it or what the kernel was doing before a packet reached that point. Kernel instrumentation, in turn, has historically involved custom modules, debug builds, or intrusive probes.

eBPF offers a programmable observation point inside the kernel while keeping the probe logic constrained and separately loaded. Operators can attach a focused program to a supported event, gather evidence during an incident, and detach it without changing application code. This is especially useful for production systems where reproducing a problem with a custom kernel or adding new application instrumentation is impractical.

The trade-off is that eBPF moves some complexity into program compatibility, permissions, data handling, and operational discipline. It is a way to observe specific kernel paths, not a shortcut around understanding what those paths mean.

How eBPF Tracing Works

A typical tracing tool turns a high-level query or compiled program into eBPF instructions and asks the kernel to load it. The kernel verifier analyzes control flow and memory access, checks that execution is bounded, and limits operations according to the program type and available helpers. Programs that fail verification are rejected. Accepted programs may be interpreted or just-in-time compiled, then attached to an eligible hook. The Linux kernel BPF documentation covers program types, maps, the loading interface, and verifier behavior.

At runtime, a matching event invokes the program. It can read permitted context, update a BPF map, or emit an event to user space. Maps are kernel-managed data structures used to retain counters, histograms, or state across invocations. Ring buffers and perf buffers provide ways to stream event records to a collector. Keeping aggregation in the kernel—for example, counting calls per process instead of sending every call—can reduce data volume.

Different hooks answer different questions. They are not interchangeable, and a probe tied to an internal implementation detail can break when the kernel changes.

Hook or approach Where it observes Useful for Important trade-off
Tracepoints Named kernel events exposed for tracing System calls, scheduler events, and other defined event paths The event must exist on the running kernel; inspect its fields and availability
kprobes and kretprobes Entry to or return from selected kernel functions Investigating a kernel function that lacks a suitable tracepoint Function names and details are less stable across kernel builds
fentry and fexit BTF-described function entry or exit Lower-overhead function tracing when supported Requires suitable kernel support and type information
uprobes and uretprobes Selected user-space executable or library functions Observing a process when source-level instrumentation is unavailable Binary, symbol, and language-runtime details can complicate portability
XDP and traffic-control hooks Network packet processing paths Early packet handling, drops, or traffic accounting Hook placement and behavior depend on mode and networking setup

For network work, an eBPF data path can reveal how packets are handled at supported points in the kernel. Cilium’s eBPF documentation describes its use of BPF in networking and related infrastructure. Protocol behavior still needs to be interpreted correctly: for example, RFC 9293 specifies TCP’s reliable, in-order byte-stream service and the sequence-number mechanisms that help detect loss.

Key Components and Concepts

Programs and hooks: The program contains the logic, while a hook determines when it runs and what context it can access. The kernel restricts helpers and data access based on the program type. A network hook, for instance, has different context and capabilities from a tracing hook.

Maps and event transport: Maps keep counters, configuration, or correlation state in the kernel. Per-CPU maps can reduce contention for high-rate counters. A ring buffer or perf buffer can deliver records to a user-space process, which can enrich, filter, export, or display them.

Loaders and front ends: Tools such as bpftrace provide a tracing language for short investigations. BCC offers libraries and tools for writing and running BPF programs, while libbpf is a common library for loading programs written with a lower-level workflow. A loader handles program loading, maps, attachment, and cleanup; it is part of the system, not merely a compiler.

BTF and CO-RE: BPF Type Format (BTF) describes kernel and program types. Compile Once – Run Everywhere (CO-RE) uses type information to adapt certain field references to the target kernel. It can ease portability, but it does not guarantee that every hook, helper, kernel configuration, or behavior exists everywhere. Check the target kernel and the tool’s compatibility requirements.

Privileges and overhead: Loading programs usually requires elevated privileges or specific capabilities, and Linux security settings, lockdown modes, container restrictions, and distribution policy can constrain access. Verified programs can still consume CPU or produce too much data. Keep probes narrow, aggregate where possible, measure their impact, and remove them when the investigation ends.

Real-World Uses

  • Finding latency below the application: Measure time spent in scheduler, filesystem, block I/O, or networking paths to distinguish application work from time waiting on kernel or device activity.
  • Investigating network drops or retransmissions: Count events at a relevant network hook and correlate them with interface, process, or service context. For a lower-level walk through packet captures and common connectivity faults, see the Linux network troubleshooting guide.
  • Understanding resource contention: Observe CPU scheduling or system-call patterns to find processes that compete for resources or issue unexpectedly frequent operations.
  • Supporting security monitoring: Linux Security Module (LSM) hooks can support policy and detection tools. This does not make eBPF a complete security boundary; production policies still need review, least privilege, and monitoring for blind spots.
  • Adding system context to service telemetry: eBPF-based tools can expose host-level behavior alongside application signals. Application spans remain useful for business operations and cross-service request context; kernel events answer a different layer of questions.

Getting Started with bpftrace

Use a development or test Linux machine first. Package names and kernel support vary by distribution. On Debian or Ubuntu, install bpftrace with the package manager; on Fedora, use dnf:

# Debian or Ubuntu
sudo apt update
sudo apt install bpftrace

# Fedora (use this instead of the Debian/Ubuntu commands above)
sudo dnf install bpftrace

Check the installation and see whether the running kernel exposes the system-call tracepoint used in this example:

bpftrace --version
sudo bpftrace -l 'tracepoint:syscalls:sys_enter_openat'

If the tracepoint is listed, count openat calls by process name. Stop the command with Ctrl+C; bpftrace prints the accumulated map when it exits:

sudo bpftrace -e 'tracepoint:syscalls:sys_enter_openat { @[comm] = count(); }'

This example records a count, not file names or file contents. The tracepoint is a Linux system-call event, so unrelated background processes may also contribute to the output. Run it briefly, on a host where you have permission, and treat process metadata as operational data that may still require privacy controls.

If the tracepoint is absent, list available related events with sudo bpftrace -l 'tracepoint:syscalls:*openat*' and consult the documentation for the target distribution and kernel. If loading is denied, check whether the command is running with the expected privileges and whether kernel lockdown, security policy, container capabilities, or system configuration restricts BPF. Do not disable host protections just to make a probe load. For a production rollout, first define the question, scope the hook and fields, test overhead, and decide how to secure and retain collected data.

Common Misconceptions

“eBPF can observe anything in any program.” A program can only attach to supported hooks and access context and helpers allowed for its type. Missing symbols, unavailable tracepoints, and kernel configuration differences limit coverage.

“Verified means free of risk.” Verification constrains what a program can do, but an expensive or overly broad probe can still add load or expose sensitive metadata. Scope, measure, and govern probes like other production instrumentation.

“eBPF replaces packet capture and distributed tracing.” eBPF can collect selected network metadata and kernel events, but it does not automatically capture full packet payloads or reconstruct application request context. Packet capture and application-level tracing remain useful for different diagnostic questions.

Changelog and Last Updated

Last updated: October 1. Initial publication.

TBO Editorial

About the Author

TBO Editorial writes about the latest updates about products and services related to Technology, Business, Finance & Lifestyle. Do get in touch if you want to share any useful article with our community.