NUMA and Memory Locality in Multicore Systems
NUMA and memory locality explain why two cores in the same server can take different amounts of time to reach the same data. The distinction matters when tuning multicore systems for databases, virtual machines, and AI inference: adding processors or memory does not guarantee that a workload will use them efficiently. This guide explains how Non-Uniform Memory Access works, how operating systems place threads and pages, and how to inspect a Linux host before changing its configuration.
What Is NUMA?
Non-Uniform Memory Access (NUMA) is a system design in which a processor can access memory attached to other processor groups, but not all memory has the same access cost. A NUMA node is a locality domain: it groups CPUs and memory that are relatively close to one another. Access to a node’s own memory is usually faster and has more available bandwidth than access to memory connected to another node.
NUMA is common in multi-socket servers, but a node is not always identical to a CPU socket. Firmware can expose multiple locality domains per socket, and some systems have nodes with memory but no CPUs. The operating system reads the platform topology and reports the nodes it can use. The Linux kernel’s NUMA memory policy documentation describes how Linux applies placement rules to process memory.
Why NUMA Exists
As core counts and memory capacity grow, it becomes difficult to connect every processor to every memory device with the same short, wide path. A single shared memory design can make processors compete for bandwidth and can constrain how many sockets or memory channels a platform supports.
NUMA scales capacity by dividing memory among locality domains. Each processor group has nearby memory controllers and channels; an inter-socket or on-board fabric carries requests to memory in another domain. A program can still address memory across the system, so this is a shared-address-space machine rather than a set of isolated computers. The cost of reaching each address, however, depends on where its physical page resides relative to the core making the request.
The trade-off is that hardware exposes more capacity and parallel bandwidth while software has to account for placement. A thread scheduled on one node may repeatedly read pages allocated on another. That extra traffic can increase latency, consume interconnect bandwidth, and compete with other requests. Performance therefore depends not just on how much RAM is installed, but on how the workload’s CPU execution and memory allocations line up with the topology.
How NUMA and Memory Locality Work
A simplified data path is:
Thread on a CPU → local memory controller → local DRAM
or, for a remote page:
Thread on a CPU → local interconnect → remote node’s memory controller → remote DRAM
The operating system sees a topology assembled from firmware and hardware information. In systems using ACPI, the ACPI specification’s System Resource Affinity Table (SRAT) describes processor and memory affinity information that helps software identify proximity domains. The OS turns that information into CPU and memory nodes, while hardware routes the actual memory transactions.
When an application requests memory, the kernel chooses physical pages according to the active memory policy. A common Linux default is local allocation: pages are allocated near the CPU that first faults them in, when capacity is available. This is often called first-touch placement. If the application initializes a large array on one thread and then shares it across many threads, the pages may end up concentrated on that thread’s node even if other nodes later do most of the work.
The scheduler also chooses which CPU runs each thread. If threads migrate frequently between nodes while their pages stay behind, memory locality can deteriorate. Operating systems can use automatic NUMA balancing to observe access patterns and adjust task or page placement, but it is not a universal fix: migration costs time, policies and workloads differ, and a workload may change faster than placement can adapt.
| Placement pattern | CPU and memory relationship | Typical effect | Useful when |
|---|---|---|---|
| Local allocation | Threads use pages from their nearby node where possible | Usually lower access latency and less cross-node traffic | A workload has enough local capacity and stable thread placement |
| Remote allocation | Threads access pages attached to another node | May add latency and use interconnect bandwidth | Local capacity is exhausted or data must be shared across nodes |
| Interleaved allocation | Pages are spread across selected nodes | Can distribute bandwidth, but each thread may access remote pages | A large shared data set benefits from aggregate bandwidth and has broad parallel access |
These are placement strategies, not fixed performance levels. Hardware generation, node distance, memory-channel population, contention, page size, and access patterns all matter. The distance values shown by Linux express relative topology costs; they are not measurements in nanoseconds.
Components and Key Concepts
- NUMA node: A set of CPUs and/or memory with similar locality. Use the OS-visible topology rather than guessing from socket count.
- Node distance: A relative indication of the cost between nodes. A higher value generally means less-local access, but it should not be read as a latency measurement.
- Memory policy: A rule for choosing nodes for new pages. Policies can prefer local memory, bind allocations to specific nodes, or interleave allocations across nodes. Linux documents process and shared-memory policies in its kernel documentation.
- CPU affinity: A limit on which CPUs can run a process or thread. Keeping a thread near its data can improve locality, but overly strict affinity can prevent load balancing.
- First touch: The placement effect where the CPU that first faults in a page influences which node supplies it under a local allocation policy. Parallel initialization can help distribute pages when it resembles later access.
- NUMA balancing: Kernel mechanisms that may move tasks or pages based on observed access. Their presence does not guarantee that every application is optimally placed.
The numactl project provides Linux commands and a library for inspecting NUMA topology and applying CPU or memory policies. The numa(7) Linux manual page gives an additional overview of the Linux NUMA model and its interfaces.
Real-World Uses
- Databases: A database may use many worker threads and a large buffer pool. If workers and hot pages are concentrated on one node, that node’s memory bandwidth can saturate while other nodes remain underused. Benchmark placement alongside query latency and throughput; pinning every thread is not automatically beneficial.
- Virtualization hosts: Hypervisors schedule vCPUs and back guest memory with host pages. A VM that spans nodes can experience different costs depending on vCPU and memory placement. Virtual NUMA (vNUMA) can expose a topology to a guest, but the guest’s view should correspond to the host’s actual allocation strategy.
- AI inference and serving: Model weights, caches, and request-processing threads compete for memory capacity and bandwidth. A larger server may help only when the serving process can use the additional nodes without adding excessive remote traffic. NUMA placement is one factor alongside batching, accelerator transfers, and model memory footprint; see the related guide to LLM serving schedulers and continuous batching.
- Scientific and high-performance computing: Parallel jobs can partition data and worker threads by node to reduce communication and make effective use of per-node memory bandwidth. The application’s decomposition and the machine’s topology must agree.
NUMA also matters when memory is attached through a fabric rather than directly to a processor. Such capacity can appear with distinct access costs and placement rules; CXL memory pooling is one example of why a system’s memory map should be understood as a topology rather than a flat pool.
Getting Started: Inspecting and Tuning a Linux Host
Start with observation before applying affinity or memory policies. Install the numactl utilities using the package manager for your distribution:
# Debian or Ubuntu
sudo apt install numactl
# Fedora or compatible distributions
sudo dnf install numactl
Check the topology and current policy:
lscpu | grep -E 'NUMA node|Socket|CPU\(s\)'
numactl --hardware
numactl --show
numactl --hardware lists available nodes, their CPUs and memory, and relative distances. CPU and node numbering varies by machine, so use the output rather than assuming that node 0 corresponds to a particular socket or that every node has the same capacity.
For a running process, inspect its NUMA allocation with:
numastat -p <PID>
less /proc/<PID>/numa_maps
Replace <PID> with the process ID. These views help identify whether memory is distributed across nodes or concentrated on a subset. Compare them with application metrics and a repeatable workload test; placement counts alone do not show whether the pages are hot or whether moving them would improve performance.
To test a process with CPUs and memory constrained to node 0, run it under an explicit policy:
numactl --cpunodebind=0 --membind=0 -- ./your-app --config app.yaml
Replace the example program and arguments, and choose a node that exists on the host. --cpunodebind restricts eligible CPUs; --membind restricts where new memory allocations can come from. If the selected node lacks enough free memory, the process may fail allocations rather than transparently using another node. To compare a distributed placement policy, test interleaving across the available nodes:
numactl --interleave=all -- ./your-app --config app.yaml
Use a representative, repeatable workload and compare throughput, tail latency, CPU utilization, memory pressure, and per-node allocation before and after the change. Test under realistic concurrency, and change one policy at a time. For production services, confirm that startup scripts, service managers, container CPU sets, and hypervisor settings do not override or conflict with the policy. Restore the original configuration if the workload regresses.
Common Misconceptions
- “NUMA nodes are always CPU sockets.” A socket can contain multiple locality domains, and firmware may report memory-only nodes. Inspect the topology the operating system actually sees.
- “Remote memory is inaccessible.” It is usually addressable; the penalty is that access has a different latency, bandwidth, and interconnect cost.
- “Binding a process to one node always makes it faster.” Binding can reduce locality variation, but it can also create CPU or memory bottlenecks and prevent the scheduler from balancing work.
- “Interleaving is the best setting for every large workload.” It distributes allocations but does not make every access local. The best policy depends on how threads share and traverse data.
- “More installed RAM fixes NUMA performance.” Additional capacity can reduce memory pressure, but it does not guarantee balanced channels, local pages, or sufficient interconnect bandwidth.
Related Articles
- Server hardware configuration and memory planning
- Hardware virtualization and virtual machines
- CXL memory pooling and fabric-attached memory
- LLM serving schedulers and continuous batching
Changelog
Published October 9. Initial explainer on NUMA topology, memory locality, Linux policies, and workload tuning.

