Skip to content

Integrate NVIDIA DRA Driver for GPUs as an operand - #1541

Closed
cdesiniotis wants to merge 9 commits into
mainfrom
dra-driver-integration
Closed

cdesiniotis wants to merge 9 commits into
mainfrom
dra-driver-integration

Conversation

@cdesiniotis

@cdesiniotis cdesiniotis commented Jul 17, 2025 •

Copy link
Copy Markdown
Contributor

Tests

K8s Cluster that does not support DRA

Cluster setup:

$ kubectl get nodes
NAME           STATUS   ROLES           AGE   VERSION
cnt-server-2   Ready    control-plane   61d   v1.31.0

Install GPU Operator with default options. Then, try enabling the DRA driver and verify the operator emits an error message indicating this is an invalid configuration:

$ kubectl patch clusterpolicies.nvidia.com/cluster-policy     --type='json'     -p='[{"op": "replace", "path": "/spec/draDriver/computeDomains/enabled", "value":true}]'
clusterpolicy.nvidia.com/cluster-policy patched

$ kubectl get clusterpolicies.nvidia.com/cluster-policy   
NAME             STATUS     AGE
cluster-policy   notReady   2025-07-18T22:24:47Z


$ kubectl get clusterpolicy -ojsonpath='{.items[0].status}{"\n"}' | jq .
{
  "conditions": [
    {
      "lastTransitionTime": "2025-07-18T22:24:47Z",
      "message": "",
      "reason": "Error",
      "status": "False",
      "type": "Ready"
    },
    {
      "lastTransitionTime": "2025-07-18T22:24:47Z",
      "message": "failed to initialize ClusterPolicy controller: ClusterPolicy validation failed: failed to validate DRA: the NVIDIA DRA driver is enabled in ClusterPolicy but Dynamic Resource Allocation is not enabled in the kubernetes cluster",
      "reason": "ReconcileFailed",
      "status": "True",
      "type": "Error"
    }
  ],
  "namespace": "gpu-operator",
  "state": "notReady"
}

If a user tries to enable the DRA driver during a helm install or upgrade, we emit an error message:

$ helm upgrade gpu-operator gpu-operator-v25.8.0-devel.tgz -n gpu-operator --set draDriver.computeDomains.enabled=true
Error: UPGRADE FAILED: execution error at (gpu-operator/templates/validation.yaml:13:4): 
Cannot enable NVIDIA DRA Driver for GPUs on a Kubernetes cluster that does not support DRA

K8s Cluster that supports DRA

Cluster setup:

$ kubectl get nodes        
NAME                   STATUS   ROLES                  AGE   VERSION
gb-nvl-043-compute03   Ready    control-plane,worker   8d    v1.32.4
gb-nvl-043-compute09   Ready    worker                 8d    v1.32.4

Install GPU Operator with default options. The DRA driver is disabled by default.

$ helm install gpu-operator gpu-operator-v25.8.0-devel.tgz  -n gpu-operator \
  --set operator.repository=ghcr.io/nvidia --set operator.version=1b983f26 \
  --set validator.repository=ghcr.io/nvidia --set validator.version=1b983f26
  
$ kubectl get clusterpolicies.nvidia.com/cluster-policy   
NAME             STATUS   AGE
cluster-policy   ready    2025-07-19T18:54:37Z

$ kubectl get pods -n gpu-operator --field-selector spec.nodeName=gb-nvl-043-compute09
NAME                                               READY   STATUS      RESTARTS   AGE
gpu-feature-discovery-2fwbj                        1/1     Running     0          2m12s
gpu-operator-node-feature-discovery-worker-4gmr4   1/1     Running     0          2m45s
nvidia-container-toolkit-daemonset-jhq5w           1/1     Running     0          2m12s
nvidia-cuda-validator-pdm7r                        0/1     Completed   0          33s
nvidia-dcgm-exporter-lxxxp                         1/1     Running     0          2m12s
nvidia-device-plugin-daemonset-8l8gh               1/1     Running     0          2m12s
nvidia-driver-daemonset-tcv7q                      1/1     Running     0          2m20s
nvidia-mig-manager-9hjr9                           1/1     Running     0          2m12s
nvidia-operator-validator-vgtjv                    1/1     Running     0          2m12s

Deploy the DRA Driver, only enable ComputeDomains:

$ kubectl patch clusterpolicies.nvidia.com/cluster-policy --type='json' -p='[{"op": "replace", "path": "/spec/draDriver/computeDomains/enabled", "value":true}]' 
clusterpolicy.nvidia.com/cluster-policy patched

$ kubectl get pods -n gpu-operator -lapp=nvidia-dra-driver-controller 
NAME                                           READY   STATUS    RESTARTS   AGE
nvidia-dra-driver-controller-b454444fc-pg4bw   1/1     Running   0          85s

$ kubectl get pods -n gpu-operator -lapp=nvidia-dra-driver-kubelet-plugin
NAME                                     READY   STATUS    RESTARTS   AGE
nvidia-dra-driver-kubelet-plugin-24755   1/1     Running   0          91s
nvidia-dra-driver-kubelet-plugin-x7m9p   1/1     Running   0          91s

$ kubectl get resourceslice
NAME                                                   NODE                   DRIVER                      POOL                   AGE
gb-nvl-043-compute03-compute-domain.nvidia.com-22tm2   gb-nvl-043-compute03   compute-domain.nvidia.com   gb-nvl-043-compute03   111s
gb-nvl-043-compute09-compute-domain.nvidia.com-87gdh   gb-nvl-043-compute09   compute-domain.nvidia.com   gb-nvl-043-compute09   111s

$ (echo -e "NODE\tLABEL\tCLIQUE"; kubectl get nodes -o json | \
    jq -r '.items[] | [.metadata.name, "nvidia.com/gpu.clique", .metadata.labels["nvidia.com/gpu.clique"]] | @tsv') | \
    column -t
NODE                  LABEL                  CLIQUE
gb-nvl-043-compute03  nvidia.com/gpu.clique  6a130f54-faaa-4b8f-847f-be44ab70f917.32766
gb-nvl-043-compute09  nvidia.com/gpu.clique  6a130f54-faaa-4b8f-847f-be44ab70f917.32766

At this point, I am able to run the simple IMEX channel injection sample and the Multi-Node MPI test from kubernetes-sigs/dra-driver-nvidia-gpu#249.

Now, let's try enabling the GPU portion of the DRA driver and verify that the operator emits an error message indicating this is an invalid configuration:

$ kubectl patch clusterpolicies.nvidia.com/cluster-policy     --type='json'     -p='[{"op": "replace", "path": "/spec/draDriver/gpus/enabled", "value":true}]'
clusterpolicy.nvidia.com/cluster-policy patched

$ kubectl get clusterpolicies.nvidia.com/cluster-policy     
NAME             STATUS     AGE
cluster-policy   notReady   2025-07-19T18:54:37Z

$ kubectl get clusterpolicy -ojsonpath='{.items[0].status}{"\n"}' | jq .
{
  "conditions": [
    {
      "lastTransitionTime": "2025-07-19T19:01:28Z",
      "message": "",
      "reason": "Error",
      "status": "False",
      "type": "Ready"
    },
    {
      "lastTransitionTime": "2025-07-19T19:01:28Z",
      "message": "failed to initialize ClusterPolicy controller: ClusterPolicy validation failed: failed to validate DRA: the device-plugin and DRA driver for GPUs cannot both be enabled in ClusterPolicy",
      "reason": "ReconcileFailed",
      "status": "True",
      "type": "Error"
    }
  ],
  "namespace": "gpu-operator",
  "state": "notReady"
}

If a user tries to enable the GPU portion of the DRA driver during a helm install or upgrade and forgets to disable the device-plugin, we emit an error message:

$ helm upgrade gpu-operator gpu-operator-v25.8.0-devel.tgz -n gpu-operator --set draDriver.gpus.enabled=true
Error: UPGRADE FAILED: execution error at (gpu-operator/templates/validation.yaml:7:4): 
NVIDIA device plugin and NVIDIA DRA Driver for GPUs cannot both be enabled

Let's disable the device plugin and then enable the GPU portion of the DRA driver:

$ kubectl patch clusterpolicies.nvidia.com/cluster-policy     --type='json'     -p='[{"op": "replace", "path": "/spec/devicePlugin/enabled", "value":false}]'           
clusterpolicy.nvidia.com/cluster-policy patched

$ kubectl patch clusterpolicies.nvidia.com/cluster-policy     --type='json'     -p='[{"op": "replace", "path": "/spec/draDriver/gpus/enabled", "value":true}]' 
clusterpolicy.nvidia.com/cluster-policy patched

$ kubectl get pods -n gpu-operator -lapp=nvidia-dra-driver-kubelet-plugin
NAME                                     READY   STATUS    RESTARTS   AGE
nvidia-dra-driver-kubelet-plugin-9zv7q   2/2     Running   0          23s
nvidia-dra-driver-kubelet-plugin-mtpvd   2/2     Running   0          20s

$  kubectl get resourceslice
NAME                                                   NODE                   DRIVER                      POOL                   AGE
gb-nvl-043-compute03-compute-domain.nvidia.com-kwz75   gb-nvl-043-compute03   compute-domain.nvidia.com   gb-nvl-043-compute03   31s
gb-nvl-043-compute03-gpu.nvidia.com-6vgqv              gb-nvl-043-compute03   gpu.nvidia.com              gb-nvl-043-compute03   30s
gb-nvl-043-compute09-compute-domain.nvidia.com-chjk5   gb-nvl-043-compute09   compute-domain.nvidia.com   gb-nvl-043-compute09   27s
gb-nvl-043-compute09-gpu.nvidia.com-th4gj              gb-nvl-043-compute09   gpu.nvidia.com              gb-nvl-043-compute09   27s

$ kubectl get clusterpolicies.nvidia.com/cluster-policy     
NAME             STATUS   AGE
cluster-policy   ready    2025-07-19T18:54:37Z

Run a sample GPU workload using DRA. The sample runs one pod with two containers. Each container shares access to the same GPU:

$ kubectl apply -f https://raw.githubusercontent.com/NVIDIA/k8s-dra-driver-gpu/refs/heads/main/demo/specs/quickstart/gpu-test2.yaml
namespace/gpu-test2 created
resourceclaimtemplate.resource.k8s.io/single-gpu created
pod/pod created

$ kubectl get pods -n gpu-test2 
NAME   READY   STATUS    RESTARTS   AGE
pod    2/2     Running   0          9s

$ kubectl logs -n gpu-test2 pod -c ctr0
GPU 0: NVIDIA GB200 (UUID: GPU-6fd2fa89-7bca-5235-b73a-0278535e1dd3)

$ kubectl logs -n gpu-test2 pod -c ctr1
GPU 0: NVIDIA GB200 (UUID: GPU-6fd2fa89-7bca-5235-b73a-0278535e1dd3)

@coveralls

coveralls commented Jul 17, 2025 •

Copy link
Copy Markdown

Coverage Status

coverage: 28.3% (+0.1%) from 28.195% — dra-driver-integration into main

@cdesiniotis
cdesiniotis force-pushed the dra-driver-integration branch 2 times, most recently from d2fb467 to e70b999 Compare July 17, 2025 19:17
Comment thread assets/state-dra-driver/0200_clusterrole.yaml Outdated
Comment thread assets/state-dra-driver/0600_configmap.yaml Outdated
- patch
- delete
- apiGroups:
- resource.k8s.io

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

For my reference, what is the k8s version that introduces these APIs?

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

According to kubernetes/enhancements#4381, 1.30 introduced this for structured parameters.

I think it may have been earlier when classical DRA was introduced.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, kubernetes/enhancements#3063.

Looks like 1.27 introduced the resource api.

Comment thread assets/state-dra-driver/0500_deployment.yaml Outdated
Comment thread assets/state-dra-driver/0700_daemonset.yaml Outdated
Comment thread assets/state-dra-driver/0500_deployment.yaml Outdated
@cdesiniotis
cdesiniotis force-pushed the dra-driver-integration branch from e70b999 to d250799 Compare July 18, 2025 01:50
Comment thread deployments/gpu-operator/templates/upgrade_crd.yaml Outdated
Comment thread assets/state-dra-driver/0200_clusterrole.yaml Outdated
// +operator-sdk:gen-csv:customresourcedefinitions.specDescriptors=true
// +operator-sdk:gen-csv:customresourcedefinitions.specDescriptors.displayName="Tolerations"
// +operator-sdk:gen-csv:customresourcedefinitions.specDescriptors.x-descriptors="urn:alm:descriptor:com.tectonic.ui:advanced,urn:alm:descriptor:io.kubernetes:Tolerations"
Tolerations []corev1.Toleration `json:"tolerations,omitempty"`

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

why Tolerations?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We add tolerations to every component managed by the gpu-operator. It's common for gpu-operator operands to run on nodes with specific taints, so we need to support setting the corresponding tolerations to the operands

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Comment thread controllers/object_controls.go Outdated
@cdesiniotis
cdesiniotis force-pushed the dra-driver-integration branch 2 times, most recently from e63ef4f to d5ec8c2 Compare July 18, 2025 23:24
Comment thread controllers/state_manager.go Outdated
@cdesiniotis
cdesiniotis force-pushed the dra-driver-integration branch from d5ec8c2 to e38e9a9 Compare July 18, 2025 23:39
Comment thread assets/state-dra-driver/0110_compute_domain_daemon-service_account.yaml Outdated
Comment thread assets/state-dra-driver/0320_compute_domain_daemon-clusterrolebinding.yaml Outdated
@cdesiniotis
cdesiniotis requested review from jgehrcke and klueska July 21, 2025 19:52
@cdesiniotis
cdesiniotis force-pushed the dra-driver-integration branch 2 times, most recently from 2f4d08a to a0da6f5 Compare September 18, 2025 06:11
@github-actions

Copy link
Copy Markdown
Contributor

This PR is stale because it has been open 90 days with no activity. This PR will be closed in 30 days unless new comments are made or the stale label is removed. To skip these checks, apply the "lifecycle/frozen" label.

@github-actions github-actions Bot added the lifecycle/stale Denotes an issue or PR has remained open with no activity and has become stale. label Feb 20, 2026
@rahulait rahulait removed the lifecycle/stale Denotes an issue or PR has remained open with no activity and has become stale. label Feb 23, 2026
@dims

dims commented Mar 3, 2026

Copy link
Copy Markdown

Is this accurate summary (so far!) @cdesiniotis ?

Integrates the NVIDIA DRA Driver for GPUs as a first-class managed operand within the GPU Operator. DRA (Dynamic Resource Allocation) is a Kubernetes-native mechanism (GA in k8s 1.32) for fine-grained, structured device sharing — an alternative to the classic device plugin model. The GPU Operator can now deploy and lifecycle-manage the DRA driver alongside existing operands.

1. New API Surface (ClusterPolicy)

Adds a draDriver section to ClusterPolicySpec with two sub-features:

  • draDriver.computeDomains — enables the compute-domain DRA driver (IMEX/NVLink fabric topology, multi-node GPU cliques)
  • draDriver.gpus — enables the GPU DRA driver (structured GPU resource advertising via ResourceSlice); mutually exclusive with the device plugin

2. New State: state-dra-driver

New manifests under assets/state-dra-driver/:

  • ServiceAccounts for driver and compute-domain daemon
  • ClusterRoles/Roles/Bindings for controller, kubelet-plugin, compute-domain daemon
  • 4 DeviceClass objects: compute-domain-daemon, compute-domain-default-channel, gpu, mig
  • Controller Deployment
  • ConfigMap for driver config
  • DaemonSet for DRA kubelet-plugin (privileged on OpenShift via dedicated RoleBinding)

3. Controller Integration

  • state_manager.go — new StateDRADriver handler registered in the state machine
  • object_controls.go — ~290 lines of DRA reconciliation (controller Deployment, kubelet-plugin DaemonSet, ConfigMap); applies standard GPU Operator transforms (image overrides, tolerations, resource limits)
  • resource_manager.go — updated for DRA cluster-level permissions

4. Validation (Two Layers)

  • Helm-time (validation.yaml): blocks install/upgrade if DRA enabled on non-DRA cluster, or device-plugin + DRA GPUs both enabled
  • Operator-time (clusterpolicy_validator.go): checks resource.k8s.io API group at startup; sets ClusterPolicy notReady with human-readable error if DRA requested but not supported

5. Helm Chart Updates

  • values.yaml: new draDriver.computeDomains.* and draDriver.gpus.* values (disabled by default)
  • clusterrole.yaml / role.yaml expanded with DRA permissions
  • clusterpolicy.yaml template updated
  • CRDs regenerated; upgrade_crd.yaml / cleanup_crd.yaml updated for ComputeDomain CRD

6. Tests

  • Unit tests in controllers/transforms_test.go (+330 lines) for DRA object transforms
  • Manual E2E documented in PR: non-DRA cluster (error path), DRA cluster with compute-domains only, DRA cluster with GPU DRA (device-plugin disabled)
Decision Rationale
DRA GPUs mutually exclusive with device-plugin Prevents double-allocation of the same GPU resources
ComputeDomains can coexist with device-plugin Compute-domain (IMEX) resources are distinct from raw GPU resources
Disabled by default DRA requires k8s ≥1.32; safe opt-in for existing users
DeviceClass apiVersion conditional Handles both v1alpha3 (k8s 1.31) and v1beta1 (k8s 1.32+)
OpenShift privileged via separate RoleBinding Follows existing GPU Operator pattern for OpenShift SCCs

@cdesiniotis
cdesiniotis force-pushed the dra-driver-integration branch from 604bd02 to f50520b Compare April 16, 2026 01:30
@cdesiniotis
cdesiniotis force-pushed the dra-driver-integration branch 3 times, most recently from e1b7387 to eb33c2b Compare April 16, 2026 01:56
Signed-off-by: Christopher Desiniotis <cdesiniotis@nvidia.com>
Signed-off-by: Christopher Desiniotis <cdesiniotis@nvidia.com>
Signed-off-by: Christopher Desiniotis <cdesiniotis@nvidia.com>
Signed-off-by: Christopher Desiniotis <cdesiniotis@nvidia.com>
Signed-off-by: Christopher Desiniotis <cdesiniotis@nvidia.com>
Signed-off-by: Christopher Desiniotis <cdesiniotis@nvidia.com>
Signed-off-by: Christopher Desiniotis <cdesiniotis@nvidia.com>
Signed-off-by: Christopher Desiniotis <cdesiniotis@nvidia.com>
Signed-off-by: Christopher Desiniotis <cdesiniotis@nvidia.com>
@cdesiniotis

Copy link
Copy Markdown
Contributor Author

Superceded by #2571

@tariq1890
tariq1890 deleted the dra-driver-integration branch August 19, 2026 20:57
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.