Kubernetes Dynamic Resource Allocation for GPUs

Updated on
9 min read

Kubernetes Dynamic Resource Allocation (DRA) gives workloads a structured way to request specialized hardware such as GPUs. It is useful to platform engineers and machine-learning teams that need to match workloads to devices by more than a simple GPU count. This explainer covers how claims, device drivers, and the scheduler work together, where DRA differs from the established device-plugin model, and what operators need before trying it.

What Is Kubernetes Dynamic Resource Allocation?

Dynamic Resource Allocation is a Kubernetes API and scheduling mechanism for assigning devices and other specialized resources to Pods. Instead of asking only for an integer number of an extended resource, a workload can reference a ResourceClaim that describes the kind of device or configuration it needs. The cluster’s DRA driver reports available devices and prepares an allocation for the workload.

DRA became stable in Kubernetes 1.35 and was first available in an earlier release. Its APIs are in the resource.k8s.io group. The Kubernetes DRA documentation describes the feature and its current limitations; the Kubernetes project maintains the platform and its API.

DRA is not a GPU driver or a GPU scheduler that runs outside Kubernetes. It is the contract between Kubernetes, a compatible device driver, and a workload. A vendor or cluster operator still installs and operates the driver that knows how to discover, allocate, configure, and expose the hardware.

Why DRA Exists

Kubernetes has long supported device plugins. A plugin can register devices such as GPUs with a node’s kubelet, after which workloads request an extended resource, for example nvidia.com/gpu: 1. This works well when the main question is how many identical devices a container needs.

Accelerator fleets are often less uniform. A cluster can contain devices with different memory capacities, models, interconnects, or specialized functions. A raw count does not express these preferences. Teams may rely on labels, separate node pools, and vendor-specific configuration to keep workloads on compatible devices. The resulting rules can be hard to maintain, and device configuration is commonly tied to a node rather than an individual workload.

DRA adds a claim-based API for describing requirements and connecting allocation to Pod placement. A driver can expose device attributes and capacity for selection, while a DeviceClass can provide a reusable category for workload authors. That makes it possible to express a request in terms of device properties rather than a vendor-specific resource count alone.

How DRA Works

The Kubernetes DRA enhancement proposal describes the design. At a high level, allocation follows this path:

  1. A compatible DRA driver discovers hardware and publishes it to Kubernetes in ResourceSlice objects. The slices describe resources, nodes, and device properties.
  2. The driver or cluster administrator makes one or more DeviceClass objects available. A class can provide a common name and selection rules for a kind of device.
  3. A workload author creates a ResourceClaim directly or defines a ResourceClaimTemplate that creates claims for Pods. The claim requests a class and, when supported, applies selectors to device attributes or capacity.
  4. A Pod references the claim. The scheduler considers both the device allocation and the other Pod placement requirements when choosing a node.
  5. The node-side driver prepares the allocated device for the container. Drivers can use the Container Device Interface (CDI) to make device access available to the workload.
  6. When the claim is no longer in use, Kubernetes and the driver release its allocation according to the driver’s implementation.

The split matters operationally. Kubernetes coordinates claims and scheduling, but the driver owns the device-specific discovery and preparation. A claim is only satisfiable if a compatible driver is installed, its devices are published, and the requested configuration is available.

Feature Device-plugin extended resources Dynamic Resource Allocation
Workload request Integer count such as nvidia.com/gpu: 1 Pod references a ResourceClaim or a claim template
Device selection Primarily quantity and node-level placement rules DeviceClass and, when exposed by the driver, selectors over device attributes and capacity
Scheduler input Extended-resource capacity reported for each node Claim requirements and device availability participate in allocation and placement
Device configuration Commonly configured for the node or plugin Can be associated with a workload’s claim when the driver supports it
Sharing Basic device-plugin resource requests do not define general sharing; vendors may add their own mechanisms A claim can be referenced for supported sharing, but the driver controls whether and how a device can be shared
Driver requirement A device plugin that registers resources with kubelet A DRA-compatible driver that publishes and prepares resources
Maturity Device plugins are a stable Kubernetes mechanism DRA is stable starting with Kubernetes 1.35

The comparison is about Kubernetes APIs, not every vendor feature. For example, GPU time-slicing can be implemented by a vendor stack without DRA. Likewise, a DRA claim does not itself divide a physical GPU into isolated partitions.

Components and Key Concepts

  • DRA driver: Vendor or operator software that discovers devices, publishes their state, handles allocation requests, and prepares allocated devices for Pods. The NVIDIA GPU Operator documentation explains the additional software components needed to operate NVIDIA GPUs on Kubernetes; verify that the driver version and deployment mode support DRA before using claims.
  • ResourceSlice: An API object through which a driver reports device inventory and properties to Kubernetes. Stale or missing slices can make a real device unavailable to the scheduler.
  • DeviceClass: A named device category and selection policy intended to separate hardware details from workload manifests. Classes and their exposed attributes depend on the installed driver and cluster configuration.
  • ResourceClaim: A request for a device configuration. A Pod can use a pre-created claim, or reference a claim template to get a separate claim for each Pod.
  • ResourceClaimTemplate: A reusable claim specification. It is useful for Jobs and Deployments whose replicas should receive independently allocated devices.
  • Scheduler and kubelet: The scheduler chooses a feasible node as part of allocation. The kubelet and driver on that node prepare the device for the container; they do not replace the driver’s hardware-specific logic.

Real-World Use Cases

Selecting GPUs by capacity or model is useful when one pool includes accelerators with different memory sizes. If the DRA driver publishes the relevant attributes and capacity, a claim can select a suitable class and device properties instead of relying only on a node label and a generic GPU count.

Running training or inference jobs can use claim templates to request a device per Pod. This ties hardware allocation to each scheduled replica and avoids manually pinning every workload to a node. The application’s framework still determines how it uses one or multiple GPUs; DRA handles allocation, not distributed training.

Sharing an allocation may help workloads that can safely use the same device. Sharing depends on the DRA driver’s capabilities and configuration, and it is not equivalent to guaranteed memory or compute isolation. Operators should validate performance, isolation, and failure behavior with the vendor’s driver before using it for multi-tenant workloads.

Supporting several accelerator types can give platform teams a more consistent way to present hardware choices to application teams. A class such as a general-purpose accelerator can hide details, while the driver and administrators keep the class definition aligned with actual devices and policy.

Getting Started with DRA

DRA requires a cluster version and a device driver that support the API. Before writing manifests, confirm that the cluster is running Kubernetes 1.35 or later, install a DRA-compatible driver, and ask the cluster operator which DeviceClass values and attributes it exposes. Kubernetes does not provide a universal GPU class name or a vendor-independent GPU driver.

The following claim template follows the pattern in the official DRA workload guide. The example driver name, class, and attribute are placeholders: replace them with values actually published in your cluster.

apiVersion: resource.k8s.io/v1
kind: ResourceClaimTemplate
metadata:
  name: gpu-claim
spec:
  spec:
    devices:
      requests:
        - name: gpu
          exactly:
            deviceClassName: example-device-class
            selectors:
              - cel:
                  expression: >-
                    device.attributes["driver.example.com"].type == "gpu" &&
                    device.capacity["driver.example.com"].memory >= quantity("64Gi")

A Pod template can reference the claim and attach it to a container:

apiVersion: batch/v1
kind: Job
metadata:
  name: gpu-job
spec:
  template:
    spec:
      restartPolicy: Never
      resourceClaims:
        - name: gpu
          resourceClaimTemplateName: gpu-claim
      containers:
        - name: worker
          image: example.invalid/worker:latest
          resources:
            claims:
              - name: gpu

The workload image is also a placeholder; use an image that can run on the target accelerator. To apply the manifests and inspect the cluster, use:

kubectl apply -f gpu-claim.yaml
kubectl apply -f gpu-job.yaml
kubectl get deviceclasses
kubectl get resourceslices
kubectl get resourceclaims -A
kubectl get pods -w

If the claim or Pod remains pending, inspect the claim and Pod events:

kubectl describe resourceclaim <claim-name>
kubectl describe pod <pod-name>
kubectl get events --sort-by=.lastTimestamp

Check first that the named class exists and that the driver is publishing matching resources. Then compare the selector with the driver’s actual attributes, available capacity, node constraints, and any taints or quotas. Once a Pod is scheduled, verify device visibility inside the container using the vendor’s diagnostic tool, such as nvidia-smi for a supported NVIDIA setup. Successful scheduling alone does not prove the device runtime or application is configured correctly.

Common Misconceptions

“DRA installs and configures the GPU software stack.” It does not. The driver, runtime integration, and any required host components still need to be installed and maintained. DRA provides the API and scheduling workflow around device allocation.

“A ResourceClaim means the GPU is automatically shared.” A claim expresses a request and can be reused where supported, but the driver determines which sharing modes are available. Sharing does not automatically provide isolation, proportional performance, or memory partitioning.

“DRA makes device plugins obsolete.” The device-plugin API remains useful for straightforward extended-resource requests and existing driver deployments. DRA is most valuable when workloads need richer device selection, claim-based allocation, or per-workload configuration; migration depends on the device vendor and cluster version.

Changelog

  • Initial publication; feature status and API guidance checked against Kubernetes documentation.
TBO Editorial

About the Author

TBO Editorial writes about the latest updates about products and services related to Technology, Business, Finance & Lifestyle. Do get in touch if you want to share any useful article with our community.