Q: Can GPU Operator + MPS split an ADA6000 into two memory partitions (30 GB + 18 GB) for vLLM? #2169
Replies: 2 comments
|
This issue is stale because it has been open 90 days with no activity. This issue will be closed in 30 days unless new comments are made or the stale label is removed. To skip these checks, apply the "lifecycle/frozen" label. |
|
The stock GPU Operator/device-plugin MPS configuration cannot express an unequal 30 GB + 18 GB split. With two MPS replicas, the plugin sets each client's memory limit to half the GPU's reported memory—roughly 24 GB here. Two pods can each request one advertised shared GPU resource, but a client needing 30 GB would exceed that limit. The implementation calculates MPS itself supports client memory limits. However, setting Assuming you mean the RTX 6000 Ada Generation, that model does not support MIG; see the supported GPU list. One option to evaluate is two time-slicing replicas, with a separately tuned This is based on the current documentation and plugin implementation; I haven't tested your models on this GPU. |
Uh oh!
There was an error while loading. Please reload this page.
Hi,
I have a requirement to run two models on a single NVIDIA ADA6000 GPU using the GPU Operator and MPS (Multi-Process Service):
I’d like to know if it’s possible to configure MPS via the GPU Operator so that the GPU can be split into these two “memory slices” (30 GB + 18 GB) to run both models simultaneously.
Thanks in advance for any guidance i am pretty new in this stuff
All reactions