Integrate NVIDIA DRA Driver for GPUs as an operand - #1541
cdesiniotis wants to merge 9 commits into
Conversation
d2fb467 to
e70b999
Compare
| - patch | ||
| - delete | ||
| - apiGroups: | ||
| - resource.k8s.io |
There was a problem hiding this comment.
For my reference, what is the k8s version that introduces these APIs?
There was a problem hiding this comment.
According to kubernetes/enhancements#4381, 1.30 introduced this for structured parameters.
I think it may have been earlier when classical DRA was introduced.
There was a problem hiding this comment.
Yes, kubernetes/enhancements#3063.
Looks like 1.27 introduced the resource api.
e70b999 to
d250799
Compare
| // +operator-sdk:gen-csv:customresourcedefinitions.specDescriptors=true | ||
| // +operator-sdk:gen-csv:customresourcedefinitions.specDescriptors.displayName="Tolerations" | ||
| // +operator-sdk:gen-csv:customresourcedefinitions.specDescriptors.x-descriptors="urn:alm:descriptor:com.tectonic.ui:advanced,urn:alm:descriptor:io.kubernetes:Tolerations" | ||
| Tolerations []corev1.Toleration `json:"tolerations,omitempty"` |
There was a problem hiding this comment.
We add tolerations to every component managed by the gpu-operator. It's common for gpu-operator operands to run on nodes with specific taints, so we need to support setting the corresponding tolerations to the operands
There was a problem hiding this comment.
To connect some dots, some vaguely related locations in code and discussion:
- https://github.com/NVIDIA/k8s-dra-driver-gpu/blob/9951be59ac1bf827d02a1f0f3dfc0cea65c7aca3/deployments/helm/nvidia-dra-driver-gpu/templates/controller.yaml#L81
- https://github.com/NVIDIA/k8s-dra-driver-gpu/blob/9951be59ac1bf827d02a1f0f3dfc0cea65c7aca3/deployments/helm/nvidia-dra-driver-gpu/values.yaml#L63
- CD controller pod: hard requirement on control-plane label? kubernetes-sigs/dra-driver-nvidia-gpu#308
e63ef4f to
d5ec8c2
Compare
d5ec8c2 to
e38e9a9
Compare
2f4d08a to
a0da6f5
Compare
|
This PR is stale because it has been open 90 days with no activity. This PR will be closed in 30 days unless new comments are made or the stale label is removed. To skip these checks, apply the "lifecycle/frozen" label. |
|
Is this accurate summary (so far!) @cdesiniotis ? Integrates the NVIDIA DRA Driver for GPUs as a first-class managed operand within the GPU Operator. DRA (Dynamic Resource Allocation) is a Kubernetes-native mechanism (GA in k8s 1.32) for fine-grained, structured device sharing — an alternative to the classic device plugin model. The GPU Operator can now deploy and lifecycle-manage the DRA driver alongside existing operands. 1. New API Surface (
|
| Decision | Rationale |
|---|---|
| DRA GPUs mutually exclusive with device-plugin | Prevents double-allocation of the same GPU resources |
| ComputeDomains can coexist with device-plugin | Compute-domain (IMEX) resources are distinct from raw GPU resources |
| Disabled by default | DRA requires k8s ≥1.32; safe opt-in for existing users |
| DeviceClass apiVersion conditional | Handles both v1alpha3 (k8s 1.31) and v1beta1 (k8s 1.32+) |
| OpenShift privileged via separate RoleBinding | Follows existing GPU Operator pattern for OpenShift SCCs |
604bd02 to
f50520b
Compare
e1b7387 to
eb33c2b
Compare
Signed-off-by: Christopher Desiniotis <cdesiniotis@nvidia.com>
Signed-off-by: Christopher Desiniotis <cdesiniotis@nvidia.com>
Signed-off-by: Christopher Desiniotis <cdesiniotis@nvidia.com>
Signed-off-by: Christopher Desiniotis <cdesiniotis@nvidia.com>
Signed-off-by: Christopher Desiniotis <cdesiniotis@nvidia.com>
Signed-off-by: Christopher Desiniotis <cdesiniotis@nvidia.com>
Signed-off-by: Christopher Desiniotis <cdesiniotis@nvidia.com>
Signed-off-by: Christopher Desiniotis <cdesiniotis@nvidia.com>
Signed-off-by: Christopher Desiniotis <cdesiniotis@nvidia.com>
eb33c2b to
3a85e9e
Compare
|
Superceded by #2571 |
Tests
K8s Cluster that does not support DRA
Cluster setup:
Install GPU Operator with default options. Then, try enabling the DRA driver and verify the operator emits an error message indicating this is an invalid configuration:
If a user tries to enable the DRA driver during a helm install or upgrade, we emit an error message:
K8s Cluster that supports DRA
Cluster setup:
Install GPU Operator with default options. The DRA driver is disabled by default.
Deploy the DRA Driver, only enable ComputeDomains:
At this point, I am able to run the simple IMEX channel injection sample and the Multi-Node MPI test from kubernetes-sigs/dra-driver-nvidia-gpu#249.
Now, let's try enabling the GPU portion of the DRA driver and verify that the operator emits an error message indicating this is an invalid configuration:
If a user tries to enable the GPU portion of the DRA driver during a helm install or upgrade and forgets to disable the device-plugin, we emit an error message:
Let's disable the device plugin and then enable the GPU portion of the DRA driver:
Run a sample GPU workload using DRA. The sample runs one pod with two containers. Each container shares access to the same GPU: