Linux io_uring Explained: Asynchronous I/O Architecture

Updated on
9 min read

When a Linux service handles many disk or network operations, repeatedly entering the kernel and waiting for each result can waste CPU time and increase latency. Linux io_uring provides a completion-based interface for submitting I/O work and collecting its results through shared queues. It is useful to systems programmers building storage tools, network servers, databases, and runtimes, but it is not an automatic speed switch: the workload, kernel, device, and way the application uses the interface all matter.

What Is Linux io_uring?

io_uring is a Linux kernel interface for asynchronous I/O. An application places operation descriptions in a submission queue (SQ), asks the kernel to process them, and later reads results from a completion queue (CQ). Each operation is represented by a submission queue entry (SQE); each finished operation produces a completion queue entry (CQE), including a result code and optional application data.

The application and kernel share memory for these queues after setup. This lets a program prepare multiple operations together instead of issuing a separate system call for every request and result. The Linux io_uring(7) manual documents the setup calls, queue behavior, supported operations, and flags. In C programs, liburing provides helpers for creating rings, preparing SQEs, submitting work, and consuming CQEs without manually encoding every operation.

The interface is available on Linux kernels that include it, but supported operations and features depend on the running kernel and configuration. Applications should check setup and operation results at runtime rather than assume every machine supports every mode.

Why io_uring Exists

Traditional blocking I/O is straightforward: call read(), wait for it to return, and process the bytes. That model can be effective for a small number of operations, but a program handling many independent requests may spend time waiting or create many threads to keep other work moving. Each syscall also crosses the user/kernel boundary, and high-frequency workloads can pay meaningful coordination and scheduling costs.

Readiness APIs such as epoll help a process wait for activity across many file descriptors, especially sockets. The application still has to issue the actual read or write and handle its result. io_uring instead lets an application submit operations and receive completion events, combining work that might otherwise require multiple interactions. It is one tool among many for performance work; the Linux performance tuning guide covers the measurement needed before deciding that I/O submission is the bottleneck.

How io_uring Works

An application first creates a ring with io_uring_setup() or a library wrapper such as io_uring_queue_init(). Setup maps the SQ and CQ into the process and returns the descriptors and memory needed to use them. The program prepares SQEs that describe work such as reading, writing, accepting a connection, or sending data. It then submits one or more entries with io_uring_enter() or a liburing helper.

The kernel processes submitted entries and posts CQEs as operations complete. A CQE carries a result—commonly a byte count or a negative error code—and an optional user_data value that the application can use to associate the result with its request. The program consumes each completion and marks it seen so the CQ can be reused. It can submit more work while earlier requests are in flight, so one thread can coordinate many operations without blocking on each one in turn.

The main interaction pattern is:

  1. Prepare one or more SQEs with operation type, file descriptor, buffers, and operation-specific parameters.
  2. Submit the entries and let the kernel make progress.
  3. Wait for or inspect CQEs.
  4. Match completions to requests, handle errors, and release or reuse application resources.
Model What the application waits for Who performs the I/O operation Typical fit
Blocking read() or write() The operation itself to return The calling thread enters the kernel and waits Simple tools and low-concurrency work
epoll readiness A descriptor becoming ready The application issues I/O after readiness Large numbers of sockets
POSIX AIO Completion notification for an AIO request Depends on the implementation and operation Portable asynchronous-I/O interfaces
io_uring A completion entry for a submitted operation The kernel processes submitted work Linux applications with many in-flight operations

POSIX defines functions such as aio_read() in its Asynchronous Input and Output specification. That standard is useful when portability is important; io_uring is Linux-specific and exposes a different set of kernel operations and controls. Neither interface guarantees that every kind of file operation will avoid blocking internally.

Key Components and Modes

Submission and completion queues are the central interface. SQEs describe requests; CQEs report their results. An application must keep buffers and other referenced resources valid until the associated operation completes. It must also consume CQEs reliably: a program that stops draining completions can run out of queue space even if submissions are otherwise correct.

Batching lets an application prepare several SQEs before a submission and process several CQEs together. This can reduce per-operation overhead, but larger batches can add latency when the program waits to fill them. Queue depth should match the workload and available resources, not be increased blindly.

Polling options change how the kernel finds work or reports completions. With IORING_SETUP_SQPOLL, a kernel thread polls the submission queue, which can reduce submission syscalls for some workloads at the cost of CPU time and setup requirements. I/O polling (IORING_SETUP_IOPOLL) is a separate mode with device and filesystem constraints; it is not a generic switch for faster reads. Default interrupt-driven processing is often the simplest starting point.

Registered resources and provided buffers can reduce repeated setup work. Registering files or buffers may help applications that reuse a known set of descriptors or memory regions. Provided-buffer rings let the kernel select from buffers supplied by the application, which is useful when receive sizes are not known in advance. These mechanisms increase lifecycle complexity: deregistration, cancellation, and shutdown must account for outstanding operations.

Operation support is not uniform. Some operations can complete immediately; others may need kernel worker threads or have restrictions based on the file type, filesystem, device, or flags. io_uring offers a common submission and completion model, not a promise that every requested operation is performed by a dedicated hardware queue or never blocks.

Real-World Use Cases

Storage engines and databases can keep many reads and writes in flight while coordinating them through completion events. File servers and media pipelines can batch work across files or connections. Network servers can use supported accept, receive, and send operations to manage connections through the same queue model. Language runtimes may expose the interface behind a higher-level asynchronous API rather than requiring application code to manage SQEs directly.

For these systems, throughput is only one concern. A server must also bound queue depth, handle partial transfers, cancel or drain requests during shutdown, and define what happens when a completion reports an error. Resource isolation still matters: polling threads and outstanding buffers consume CPU and memory within the service’s Linux cgroup resource limits. At the application layer, kernel I/O is only one part of an asynchronous real-time service.

Getting Started with liburing

Use a Linux development environment with a kernel that supports io_uring, a C compiler, and the liburing development package. On Debian or Ubuntu, install build-essential and liburing-dev; on Fedora, install gcc and liburing-devel. Package names and available liburing versions vary by distribution.

The following small program submits one read request and prints the bytes returned. Save it as read-once.c:

#include <errno.h>
#include <fcntl.h>
#include <liburing.h>
#include <stdio.h>
#include <string.h>
#include <unistd.h>

int main(int argc, char **argv) {
    if (argc != 2) {
        fprintf(stderr, "Usage: %s FILE\n", argv[0]);
        return 2;
    }

    int fd = open(argv[1], O_RDONLY | O_CLOEXEC);
    if (fd < 0) {
        perror("open");
        return 1;
    }

    struct io_uring ring;
    int ret = io_uring_queue_init(8, &ring, 0);
    if (ret < 0) {
        fprintf(stderr, "io_uring_queue_init: %s\n", strerror(-ret));
        close(fd);
        return 1;
    }

    char buffer[4096];
    struct io_uring_sqe *sqe = io_uring_get_sqe(&ring);
    if (sqe == NULL) {
        fprintf(stderr, "No submission queue entry available\n");
        io_uring_queue_exit(&ring);
        close(fd);
        return 1;
    }

    io_uring_prep_read(sqe, fd, buffer, sizeof(buffer), 0);
    ret = io_uring_submit(&ring);
    if (ret < 0) {
        fprintf(stderr, "io_uring_submit: %s\n", strerror(-ret));
        io_uring_queue_exit(&ring);
        close(fd);
        return 1;
    }
    if (ret == 0) {
        fprintf(stderr, "No I/O requests were submitted\n");
        io_uring_queue_exit(&ring);
        close(fd);
        return 1;
    }

    struct io_uring_cqe *cqe;
    ret = io_uring_wait_cqe(&ring, &cqe);
    if (ret < 0) {
        fprintf(stderr, "io_uring_wait_cqe: %s\n", strerror(-ret));
        io_uring_queue_exit(&ring);
        close(fd);
        return 1;
    }

    int result = cqe->res;
    io_uring_cqe_seen(&ring, cqe);
    io_uring_queue_exit(&ring);
    close(fd);

    if (result < 0) {
        errno = -result;
        perror("read");
        return 1;
    }
    if (result > 0 && fwrite(buffer, 1, (size_t)result, stdout) != (size_t)result) {
        perror("write");
        return 1;
    }
    return 0;
}

Compile and run it with liburing’s compiler and linker flags:

cc -O2 -Wall -Wextra read-once.c -o read-once $(pkg-config --cflags --libs liburing)
./read-once /etc/hostname
uname -r
pkg-config --modversion liburing

The example performs one read at offset zero, so it is a demonstration of submission and completion rather than a complete file-copy loop. Production code must handle partial reads, offsets, multiple in-flight operations, cancellation, and buffer ownership. If queue setup fails with a permission error, the environment may restrict io_uring through its security policy; check the host or container configuration rather than disabling security controls blindly.

Benchmark the application on representative files and devices, compare it with a simpler implementation, and record latency as well as throughput. Start with ordinary submission and completion, then test batching or registered resources only if profiling shows syscall or coordination overhead is material. For CPU or memory limits in containers, inspect the applicable cgroup controls as part of the same measurement.

Common Misconceptions

  • “io_uring makes every file operation nonblocking.” The API is asynchronous at the request/completion boundary, but some operations or file types may use worker threads or encounter blocking behavior inside the kernel.
  • “Fewer syscalls always means faster software.” Queue setup, synchronization, memory management, and completion processing have costs. A short-lived or low-volume program may be faster and easier to maintain with ordinary system calls.
  • “Polling is always lower latency.” Polling can avoid some wakeups or submissions but consumes CPU and may require specific hardware or permissions. Measure it under the service’s actual load and resource limits.
  • “io_uring replaces epoll everywhere.” They use different models and can complement each other. Existing readiness-based code may not benefit enough to justify a rewrite.
TBO Editorial

About the Author

TBO Editorial writes about the latest updates about products and services related to Technology, Business, Finance & Lifestyle. Do get in touch if you want to share any useful article with our community.