NVIDIA GPUDirect Storage: Architecture Explained

Updated on
8 min read

GPU Direct Storage, commonly called GPUDirect Storage (GDS), is a way to move data between supported storage and NVIDIA GPU memory without routing the payload through an intermediate CPU buffer. It matters to machine-learning engineers and infrastructure teams when data loading limits accelerator throughput. This explainer follows the I/O path, explains which parts still depend on the host, and shows how to check whether a system can use GDS.

What Is GPU Direct Storage?

GPUDirect Storage is NVIDIA’s software and platform capability for direct data transfers between a storage system and GPU memory. Traditional applications often read a file into a host-memory buffer and then copy that data to the GPU. Where the hardware, drivers, and storage path support GDS, the data can instead move by direct memory access (DMA) between storage and GPU memory, avoiding that intermediate host-memory bounce buffer.

GDS is not a new storage protocol, a GPU, or a general operating-system setting that automatically accelerates every file read. An application must use NVIDIA’s cuFile APIs or a framework that integrates them, and the complete I/O path must be supported. NVIDIA’s GPUDirect Storage documentation describes the architecture, supported environments, and APIs; the NVIDIA GDS project page provides a product overview.

Why It Exists

AI training and other data-intensive GPU workloads repeatedly load large datasets, checkpoints, and model files. In a conventional path, storage transfers data to host memory, after which the CPU or a copy engine moves it across PCIe to GPU memory. This requires host buffers and can consume CPU cycles and memory bandwidth. When a workload is I/O-bound, the GPU may wait for data even though the storage device and accelerator each have unused capacity.

GDS targets the data movement portion of that bottleneck. Removing an extra payload copy can reduce host-memory traffic and CPU involvement, and can improve throughput or latency for suitable workloads. It does not make the storage device faster, remove filesystem work, or guarantee a faster application. If preprocessing, random access, network congestion, or GPU compute is the actual bottleneck, a direct transfer alone may not change end-to-end performance.

The underlying storage interface remains important. For example, NVM Express specifications define the NVMe command set and transport specifications; they do not define GDS. NVIDIA implements the direct GPU data path on top of supported storage and software stacks, which can include local NVMe and selected networked storage configurations.

How the Data Path Works

With a conventional read, the application requests file data, the operating system and filesystem service the request, and the payload lands in host memory. The application then copies it into a GPU buffer. That route is broadly compatible but adds a staging step.

With GDS, an application registers or supplies a GPU buffer and uses cuFile operations such as cuFileRead or cuFileWrite. The cuFile library coordinates with the filesystem, NVIDIA’s kernel components where required, storage drivers, and the GPU. On a supported topology, a DMA-capable device transfers the payload directly to or from GPU memory over the platform interconnect. The NVIDIA CUDA Toolkit GDS documentation describes these application interfaces.

The CPU does not disappear from the operation. It still runs the application, submits I/O, handles filesystem and metadata work, and manages parts of the transfer lifecycle. “Direct” describes the payload’s data path, not an absence of CPU work or an automatic bypass of every software layer.

Path Payload route Host-memory staging Main trade-off
Conventional I/O plus GPU copy Storage → host buffer → GPU buffer Yes Broad compatibility, but includes an additional data copy
GDS direct path Supported storage → GPU buffer Avoided for the supported transfer Less host-memory traffic; requires a supported stack and topology
GDS compatibility path Storage → host buffer → GPU buffer Yes Lets cuFile-based applications run where a direct path is unavailable, but may not provide direct-path performance

GDS can use a compatibility path when direct I/O requirements are not met, depending on configuration. In that case, an application using cuFile may still work, but the transfer may fall back to host-memory staging. That is why seeing the library installed is not proof that a benchmark used the direct path.

Components and Key Concepts

  • cuFile: The user-space API through which an application or framework submits file I/O involving GPU buffers. The library provides synchronous and asynchronous interfaces; the right choice depends on the application’s concurrency and buffering model.
  • Kernel and storage integration: The filesystem, kernel, and storage drivers must support the transfer path. Depending on the platform and release, GDS may use NVIDIA’s nvidia-fs kernel module or supported Linux peer-to-peer DMA facilities. Installation requirements change, so check the NVIDIA compatibility and installation documentation for the deployed release.
  • GPU memory and DMA: The target buffer must be accessible for the transfer. Registration, alignment, transfer size, and buffer reuse can affect overhead and throughput.
  • PCIe topology: Device placement and PCIe routing matter. A GPU and storage device behind a topology that cannot support peer-to-peer transfers may force a different route. NVIDIA’s best-practices guide discusses topology, Access Control Services (ACS), and IOMMU considerations.
  • Direct and compatibility modes: Direct mode avoids the CPU bounce buffer for supported I/O. Compatibility mode prioritizes operability and may use host staging, so it should not be treated as a performance equivalent.

GDS is also different from GPU-to-GPU peer transfers and from RDMA as a standalone network feature. Those technologies can share hardware capabilities, but GDS specifically concerns file or storage I/O into GPU memory through supported software and devices.

Real-World Use Cases

Training large models can benefit when dataset reads keep GPUs idle. A direct storage path may reduce staging overhead, particularly when sustained reads are large enough to amortize setup costs. Data decoding, augmentation, and sampling can still dominate, so measure the complete input pipeline.

Inference and retrieval systems may read model weights, embedding data, or other large files into accelerator memory. The benefit depends on access patterns and caching: a workload serving data already resident in memory may gain little from optimizing storage reads.

Checkpointing and data pipelines can use the write path as well as reads. Saving checkpoints directly from GPU buffers may reduce host-memory copies, but durability, filesystem semantics, and application consistency remain the responsibility of the surrounding system.

GDS can be part of a containerized or multi-GPU platform, but it addresses I/O after a workload has access to a GPU. It does not allocate accelerators; Kubernetes operators can manage that layer separately with mechanisms such as Kubernetes Dynamic Resource Allocation for GPUs. For the broader compute and data-loading context, see multi-GPU machine-learning setup.

Getting Started and Verifying a Path

GDS is primarily deployed on Linux systems with NVIDIA GPUs. Start by identifying the GPU, driver, CUDA Toolkit, storage device or filesystem, kernel, and PCIe topology. Compare the complete combination against the current NVIDIA support matrix; a supported GPU alone is not enough. Install the driver, CUDA Toolkit, GDS components, and any filesystem-specific requirements using the instructions for that platform. Package names and whether nvidia-fs is installed separately can vary by release, so a universal apt install command would be misleading.

The following checks assume the GDS tools directory is on PATH:

# Confirm the NVIDIA driver can see the GPU
nvidia-smi

# Check platform and GDS capabilities
gdscheck.py -p

# Inspect the installed I/O benchmark's supported options
gdsio -h

Use the bundled gdsio utility to benchmark supported reads and writes. Compare it with a conventional baseline using the same storage, file, block size, queue depth, and concurrency. Measure throughput, latency, CPU use, and GPU idle time; include warm and cold-cache runs where relevant. A result is useful only if the test confirms which path it exercised rather than silently measuring compatibility mode.

The cuFile configuration is typically stored in /etc/cufile.json. Inspect the current settings before changing them, and adjust them only after establishing a baseline and consulting the parameter documentation. For application integration, use cuFile directly or a framework such as KVikIO that exposes GDS capabilities. The API is not a drop-in replacement for every POSIX read() call; the application must pass supported file handles and GPU buffers through the integration.

For deployments with NVMe over Fabrics, validate the network adapter, filesystem, and storage topology in addition to the local GPU path. More generally, block-storage performance fundamentals help separate a storage bottleneck from a GPU-transfer bottleneck.

Common Misconceptions

“GDS eliminates the CPU from storage I/O.” It avoids a host-memory bounce buffer on the direct payload path; the CPU still submits and coordinates work and handles other software responsibilities.

“Enabling GDS makes every workload faster.” Benefits depend on the amount and pattern of I/O, storage performance, software support, and topology. Small or compute-bound workloads may show little improvement, and tuning can change results.

“A successful cuFile read proves direct storage-to-GPU DMA.” The API can use a compatibility path when a direct path is unavailable. Check platform diagnostics and benchmark behavior instead of inferring the route from successful completion alone.

Changelog

  • Initial publication; architecture and verification guidance checked against NVIDIA’s GDS documentation.
TBO Editorial

About the Author

TBO Editorial writes about the latest updates about products and services related to Technology, Business, Finance & Lifestyle. Do get in touch if you want to share any useful article with our community.