RDMA and RoCEv2 Networking Explained
Remote Direct Memory Access (RDMA) is used in storage and compute clusters where moving data through the ordinary operating-system network path can consume too much CPU or add too much latency. Infrastructure engineers working with AI training, high-performance computing (HPC), or NVMe storage need to understand both the RDMA operation and the fabric carrying it. This guide explains the verbs and memory protections behind RDMA, compares InfiniBand with RoCEv2, and covers how to verify a Linux setup without treating Ethernet as automatically lossless.
What Is RDMA and RoCEv2?
RDMA is a capability that lets one machine transfer data to or from memory on another machine with the network adapters handling the data movement. Instead of asking the remote CPU to copy every byte through its normal socket path, an application establishes a connection, registers memory, and gives a peer tightly scoped permission to access a region. The remote CPU still participates in setup and application logic; it is the data transfer that can avoid per-byte processing by both CPUs.
RDMA describes memory operations, not one particular cable or network protocol. InfiniBand provides its own fabric for RDMA. RoCE, or RDMA over Converged Ethernet, carries RDMA traffic over Ethernet. RoCEv2 encapsulates the traffic in UDP/IP, so IP routers can forward it between subnets when the network and endpoints are configured for it. RoCEv1 is a distinct, Layer 2 form and is generally limited to a local Ethernet broadcast domain.
The InfiniBand Trade Association (IBTA) publishes specifications covering InfiniBand and RoCE. The IBTA specification resources are the reference point for those fabric standards. The Linux kernel’s userspace verbs documentation explains how applications interact with the Linux RDMA subsystem.
The Problem RDMA Solves
With conventional TCP networking, an application sends and receives through sockets. The operating system processes packet headers, protocol state, buffers, and interrupts or polling work. This is a flexible, widely compatible approach, but at high message rates it can use significant CPU time and add latency. The problem becomes more noticeable when nodes exchange small messages frequently or sustain high bandwidth.
RDMA-capable network interface cards (RNICs) move much of that work into hardware. An application can post work requests to a queue and let the RNIC transfer payloads directly between registered memory regions. That can reduce CPU overhead and make latency more predictable, which is valuable when dozens or hundreds of servers communicate concurrently.
RDMA is not automatically the right choice for every system. It adds requirements around capable adapters, drivers, compatible libraries, memory registration, and fabric operations. A conventional TCP transport can be easier to deploy and troubleshoot. The tradeoff is similar in NVMe over Fabrics, where an RDMA transport can reduce overhead but needs a compatible network and storage stack.
How the RDMA Data Path Works
Most applications use a verbs API to create and manage RDMA resources. The Linux rdma-core project provides userspace libraries and utilities for this ecosystem. A typical connection has these stages:
- The application selects an RDMA device and creates a protection domain, which groups related resources.
- It registers a memory region (MR). Registration pins or maps the memory for the adapter and associates access permissions and keys with it.
- The application creates a queue pair (QP), exchanges connection details with its peer, and moves the QP into a usable state.
- It posts work requests to a send or receive queue. For an RDMA write or read, the request identifies the remote region and uses a remote key (rkey) granted by the peer.
- The RNIC performs the transfer and reports completion through a completion queue (CQ). The application can then reuse buffers or process the result.
An RDMA write lets one side place data into a permitted remote buffer; an RDMA read lets it retrieve data from that buffer. Send/receive operations instead involve a matching receive buffer posted by the remote application. RDMA atomics are also available on supported hardware, with specific limits and ordering rules.
The remote memory key is a capability, not an address that should be exposed without controls. Applications must scope registrations, keys, and lifetimes carefully: a peer can access only the registered range and permitted operations, but a leaked or overly broad key still creates risk. Connection setup and completion handling also consume CPU cycles. “Direct” describes the data movement, not a CPU-free system.
| Feature | InfiniBand | RoCEv2 | iWARP |
|---|---|---|---|
| Fabric | Native InfiniBand links and switches | Ethernet with UDP/IP encapsulation | Standard IP network using TCP |
| Routing | InfiniBand subnet routing | IP routing is possible | IP routing is possible |
| Endpoint | InfiniBand-capable adapter | Ethernet RNIC with RoCE support | Adapter and stack with iWARP support |
| Congestion and loss | Fabric-specific flow and congestion mechanisms | Must be planned for the Ethernet path and RNIC behavior | TCP congestion control applies |
| Typical fit | Dedicated HPC or storage fabric | RDMA on an Ethernet data-center fabric | RDMA where a TCP-based fabric is preferred |
| Main operational cost | Dedicated fabric skills and equipment | Coordinated Ethernet, QoS, and RNIC configuration | Hardware and software availability can limit choices |
iWARP is a separate RDMA transport, not another name for RoCE: RFC 5040 defines the RDMA Protocol (RDMAP) over the Direct Data Placement Protocol, which is used with iWARP. That distinction matters when comparing standards and troubleshooting an adapter stack.
Components and Key Concepts
- RNIC: The adapter and firmware that execute RDMA operations. Capabilities vary by device, driver, and firmware; support for an Ethernet link alone does not imply RoCE support.
- Verbs: The programming interface for creating resources and posting operations. Not every application must call verbs directly; storage drivers and communication libraries can use them underneath.
- QP and CQ: A queue pair carries work requests for a connection, while a completion queue records finished work and errors. Applications must size and service these queues for their workload.
- Memory registration and keys: Registration makes memory available to the adapter with specified permissions. Local keys (lkey) and remote keys (rkey) help enforce access boundaries.
- Fabric and congestion control: The switches, links, routing, and endpoint settings determine whether traffic reaches peers and how the network responds under contention. With RoCEv2, UDP/IP makes routing possible but does not configure the Ethernet network for reliable high-load operation.
Some deployments tune RoCEv2 Ethernet for low loss using Priority Flow Control (PFC), Explicit Congestion Notification (ECN), and an RNIC congestion-control algorithm. PFC pauses traffic for a selected priority on a link; it is not end-to-end congestion control and can spread congestion or contribute to deadlock if designed poorly. ECN marks packets as queues build, allowing endpoints to reduce sending rates when their RNIC algorithms support and use those signals. The right combination depends on the adapter and switch implementation. RoCEv2 does not, by itself, require every Ethernet class to be made lossless.
Real-World Use Cases
Distributed storage uses RDMA to reduce transport overhead between clients and storage targets. NVMe-oF can carry NVMe commands over RDMA, commonly over InfiniBand or RoCE, while the filesystem and storage stack continue to handle access and durability.
AI and HPC clusters exchange model data, gradients, collective-operation payloads, or simulation state between nodes. Faster communication can help when a workload is network-bound, but it cannot compensate for slow storage, inefficient algorithms, or poor scheduling. RDMA is also distinct from GPU Direct Storage: that technology concerns supported paths into GPU memory, while RDMA describes network memory transfers. See the GPUDirect Storage architecture for the storage-to-GPU path.
Virtualization and cloud platforms may expose RDMA to selected guests or workloads for latency-sensitive applications. This requires deliberate device assignment, isolation, and operational monitoring; making a physical RNIC available to a virtual machine does not automatically provide safe sharing or correct fabric settings.
Getting Started: Inspect and Test a Linux RDMA Setup
Begin with a compatible adapter, supported kernel driver and firmware, and a peer on the intended fabric. On Debian or Ubuntu, the user tools and a basic performance test can be installed with:
sudo apt-get update
sudo apt-get install rdma-core ibverbs-utils perftest
Package names and driver setup vary across distributions and hardware. The upstream rdma-core project documents the userspace components; the adapter vendor’s installation instructions remain necessary for its driver and firmware. These packages do not enable RoCE or configure switches.
Inspect the local device and link before testing:
# Show RDMA devices and links known to the kernel
rdma dev show
rdma link show
# Show verbs devices and their capabilities
ibv_devices
ibv_devinfo
# Check Ethernet link counters (replace the interface name)
ip -s link show dev enp5s0
ethtool -S enp5s0
rdma link show should report a usable state, and ibv_devinfo should list the expected adapter. For RoCEv2, also confirm the IP route, VLAN, address, MTU, and quality-of-service markings along the complete path. Vendor counter names differ, so inspect the adapter and switch documentation for dropped packets, retransmissions, ECN marks, PFC pauses, and queue congestion.
With two authorized test hosts on the configured fabric, run a bandwidth test from the server side and then connect from the client:
# Host A: wait for a peer
ib_write_bw
# Host B: replace with Host A's reachable address
ib_write_bw 192.0.2.10
Use a lab or approved maintenance window; a throughput test can load the link. Compare RDMA results with a TCP baseline such as iperf3, keeping message sizes, duration, and concurrency comparable. Check error output and counters rather than treating a single bandwidth number as proof that congestion handling is healthy.
For production RoCEv2, validate settings end to end: consistent VLAN and priority mapping, compatible MTUs, intended ECN thresholds and endpoint response, and narrowly scoped PFC where the design requires it. Test mixed traffic and failure conditions, not only an idle link. A common failure pattern is a configuration that passes a small test but stalls when multiple senders contend for a switch queue.
Common Misconceptions
“RDMA means the CPU is not involved.” The RNIC can move payloads without a remote CPU copying each byte, but software still creates connections, registers memory, posts work, and handles completions and failures.
“RoCEv2 is lossless because it uses RDMA.” RoCEv2 is an encapsulation, not a guarantee about switch queues, packet loss, or congestion. PFC and ECN are configuration tools with tradeoffs, not automatic properties of Ethernet RDMA.
“RDMA always beats TCP.” It can reduce CPU overhead and latency for suitable workloads, but hardware, software, queue setup, message size, and fabric behavior determine results. For simpler or less demanding systems, TCP can be more practical and perform well.
Related Articles
- NVMe over Fabrics (NVMe-oF) Networking explains how RDMA fits into networked storage.
- GPUDirect Storage: Architecture Explained distinguishes storage-to-GPU transfers from network RDMA.
- High-Performance Computing Architecture covers the cluster context where low-latency interconnects are used.
- Network Performance Optimization explains latency, throughput, and network troubleshooting fundamentals.
Changelog
- Initial publication.

