Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
9 changes: 9 additions & 0 deletions container-images/ramalama/Containerfile
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,15 @@ RUN ./build_llama.sh ramalama

FROM quay.io/fedora/fedora:44

# NVIDIA_DRIVER_CAPABILITIES gates which driver libraries the NVIDIA container
# toolkit injects. The Vulkan ICD only comes in under the "graphics" capability,
# and the legacy nvidia-container-runtime hook that backs "docker --gpus all"
# defaults to "compute,utility" when the variable is unset, so the ICD is
# silently absent under Docker. The hook reads the variable from the image, so
# setting it here covers "ramalama run" and the generated quadlet, kube, compose
# and stack artifacts alike. CDI, which the podman path uses, ignores it.
ENV NVIDIA_DRIVER_CAPABILITIES=compute,utility,graphics

RUN --mount=type=bind,from=builder,source=/tmp/install,target=/tmp/install,relabel=private \
cp -a /tmp/install/bin/ /usr/ && \
cp -a /tmp/install/lib64/*.so* /usr/lib64/
Expand Down
5 changes: 5 additions & 0 deletions container-images/scripts/build_llama.sh
Original file line number Diff line number Diff line change
Expand Up @@ -142,6 +142,11 @@ dnf_install_runtime_deps() {
if [ "$uname_m" = "x86_64" ] || [ "$uname_m" = "aarch64" ]; then
dnf copr enable -y slp/mesa-libkrun-vulkan
runtime_pkgs+=(vulkan-loader vulkan-tools "mesa-vulkan-drivers-$MESA_VULKAN_VERSION")
# NVIDIA's Vulkan ICD (libGLX_nvidia.so.0, injected by the container
# toolkit) needs libXext and libEGL.so.1 present to initialize.
# libglvnd-egl pulls in mesa-libEGL, so pin it to the copr version to
# avoid dragging the rest of mesa off MESA_VULKAN_VERSION.
runtime_pkgs+=(libXext libglvnd-egl "mesa-libEGL-$MESA_VULKAN_VERSION")
else
runtime_pkgs+=(openblas)
fi
Expand Down
4 changes: 2 additions & 2 deletions docs/options/backend.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,7 +10,7 @@ Available backends depend on the detected GPU hardware.

**auto** (default): Automatically selects the preferred backend based on your GPU:
- **AMD GPUs**: vulkan (Linux/macOS) or rocm (Windows)
- **NVIDIA GPUs**: cuda
- **NVIDIA GPUs**: cuda; vulkan available as explicit option
- **Intel GPUs**: vulkan (Linux/macOS) or sycl (Windows); openvino available as explicit option
- **Ascend NPUs**: cann
- **MUSA GPUs**: musa
Expand All @@ -21,7 +21,7 @@ Available backends depend on the detected GPU hardware.
- On **Windows**, vulkan is not supported on WSL2, so vendor-specific backends (rocm, sycl) are preferred

**Explicit backend selection**:
- **vulkan**: Use Vulkan-based inference (compatible with AMD, Intel, and CPU)
- **vulkan**: Use Vulkan-based inference (compatible with AMD, NVIDIA, Intel, and CPU)
- **rocm**: Use AMD ROCm backend (AMD GPUs only)
- **cuda**: Use NVIDIA CUDA backend (NVIDIA GPUs only)
- **sycl**: Use Intel SYCL/oneAPI backend (Intel GPUs only)
Expand Down
4 changes: 2 additions & 2 deletions docs/ramalama-bench.1.md
Original file line number Diff line number Diff line change
Expand Up @@ -46,7 +46,7 @@ Available backends depend on the detected GPU hardware.

**auto** (default): Automatically selects the preferred backend based on your GPU:
- **AMD GPUs**: vulkan (Linux/macOS) or rocm (Windows)
- **NVIDIA GPUs**: cuda
- **NVIDIA GPUs**: cuda; vulkan available as explicit option
- **Intel GPUs**: vulkan (Linux/macOS) or sycl (Windows); openvino available as explicit option
- **Ascend NPUs**: cann
- **MUSA GPUs**: musa
Expand All @@ -57,7 +57,7 @@ Available backends depend on the detected GPU hardware.
- On **Windows**, vulkan is not supported on WSL2, so vendor-specific backends (rocm, sycl) are preferred

**Explicit backend selection**:
- **vulkan**: Use Vulkan-based inference (compatible with AMD, Intel, and CPU)
- **vulkan**: Use Vulkan-based inference (compatible with AMD, NVIDIA, Intel, and CPU)
- **rocm**: Use AMD ROCm backend (AMD GPUs only)
- **cuda**: Use NVIDIA CUDA backend (NVIDIA GPUs only)
- **sycl**: Use Intel SYCL/oneAPI backend (Intel GPUs only)
Expand Down
4 changes: 2 additions & 2 deletions docs/ramalama-perplexity.1.md
Original file line number Diff line number Diff line change
Expand Up @@ -46,7 +46,7 @@ Available backends depend on the detected GPU hardware.

**auto** (default): Automatically selects the preferred backend based on your GPU:
- **AMD GPUs**: vulkan (Linux/macOS) or rocm (Windows)
- **NVIDIA GPUs**: cuda
- **NVIDIA GPUs**: cuda; vulkan available as explicit option
- **Intel GPUs**: vulkan (Linux/macOS) or sycl (Windows); openvino available as explicit option
- **Ascend NPUs**: cann
- **MUSA GPUs**: musa
Expand All @@ -57,7 +57,7 @@ Available backends depend on the detected GPU hardware.
- On **Windows**, vulkan is not supported on WSL2, so vendor-specific backends (rocm, sycl) are preferred

**Explicit backend selection**:
- **vulkan**: Use Vulkan-based inference (compatible with AMD, Intel, and CPU)
- **vulkan**: Use Vulkan-based inference (compatible with AMD, NVIDIA, Intel, and CPU)
- **rocm**: Use AMD ROCm backend (AMD GPUs only)
- **cuda**: Use NVIDIA CUDA backend (NVIDIA GPUs only)
- **sycl**: Use Intel SYCL/oneAPI backend (Intel GPUs only)
Expand Down
4 changes: 2 additions & 2 deletions docs/ramalama-run.1.md
Original file line number Diff line number Diff line change
Expand Up @@ -58,7 +58,7 @@ Available backends depend on the detected GPU hardware.

**auto** (default): Automatically selects the preferred backend based on your GPU:
- **AMD GPUs**: vulkan (Linux/macOS) or rocm (Windows)
- **NVIDIA GPUs**: cuda
- **NVIDIA GPUs**: cuda; vulkan available as explicit option
- **Intel GPUs**: vulkan (Linux/macOS) or sycl (Windows); openvino available as explicit option
- **Ascend NPUs**: cann
- **MUSA GPUs**: musa
Expand All @@ -69,7 +69,7 @@ Available backends depend on the detected GPU hardware.
- On **Windows**, vulkan is not supported on WSL2, so vendor-specific backends (rocm, sycl) are preferred

**Explicit backend selection**:
- **vulkan**: Use Vulkan-based inference (compatible with AMD, Intel, and CPU)
- **vulkan**: Use Vulkan-based inference (compatible with AMD, NVIDIA, Intel, and CPU)
- **rocm**: Use AMD ROCm backend (AMD GPUs only)
- **cuda**: Use NVIDIA CUDA backend (NVIDIA GPUs only)
- **sycl**: Use Intel SYCL/oneAPI backend (Intel GPUs only)
Expand Down
4 changes: 2 additions & 2 deletions docs/ramalama-sandbox-goose.1.md
Original file line number Diff line number Diff line change
Expand Up @@ -54,7 +54,7 @@ Available backends depend on the detected GPU hardware.

**auto** (default): Automatically selects the preferred backend based on your GPU:
- **AMD GPUs**: vulkan (Linux/macOS) or rocm (Windows)
- **NVIDIA GPUs**: cuda
- **NVIDIA GPUs**: cuda; vulkan available as explicit option
- **Intel GPUs**: vulkan (Linux/macOS) or sycl (Windows); openvino available as explicit option
- **Ascend NPUs**: cann
- **MUSA GPUs**: musa
Expand All @@ -65,7 +65,7 @@ Available backends depend on the detected GPU hardware.
- On **Windows**, vulkan is not supported on WSL2, so vendor-specific backends (rocm, sycl) are preferred

**Explicit backend selection**:
- **vulkan**: Use Vulkan-based inference (compatible with AMD, Intel, and CPU)
- **vulkan**: Use Vulkan-based inference (compatible with AMD, NVIDIA, Intel, and CPU)
- **rocm**: Use AMD ROCm backend (AMD GPUs only)
- **cuda**: Use NVIDIA CUDA backend (NVIDIA GPUs only)
- **sycl**: Use Intel SYCL/oneAPI backend (Intel GPUs only)
Expand Down
4 changes: 2 additions & 2 deletions docs/ramalama-sandbox-opencode.1.md
Original file line number Diff line number Diff line change
Expand Up @@ -54,7 +54,7 @@ Available backends depend on the detected GPU hardware.

**auto** (default): Automatically selects the preferred backend based on your GPU:
- **AMD GPUs**: vulkan (Linux/macOS) or rocm (Windows)
- **NVIDIA GPUs**: cuda
- **NVIDIA GPUs**: cuda; vulkan available as explicit option
- **Intel GPUs**: vulkan (Linux/macOS) or sycl (Windows); openvino available as explicit option
- **Ascend NPUs**: cann
- **MUSA GPUs**: musa
Expand All @@ -65,7 +65,7 @@ Available backends depend on the detected GPU hardware.
- On **Windows**, vulkan is not supported on WSL2, so vendor-specific backends (rocm, sycl) are preferred

**Explicit backend selection**:
- **vulkan**: Use Vulkan-based inference (compatible with AMD, Intel, and CPU)
- **vulkan**: Use Vulkan-based inference (compatible with AMD, NVIDIA, Intel, and CPU)
- **rocm**: Use AMD ROCm backend (AMD GPUs only)
- **cuda**: Use NVIDIA CUDA backend (NVIDIA GPUs only)
- **sycl**: Use Intel SYCL/oneAPI backend (Intel GPUs only)
Expand Down
4 changes: 2 additions & 2 deletions docs/ramalama-sandbox-pi.1.md
Original file line number Diff line number Diff line change
Expand Up @@ -58,7 +58,7 @@ Available backends depend on the detected GPU hardware.

**auto** (default): Automatically selects the preferred backend based on your GPU:
- **AMD GPUs**: vulkan (Linux/macOS) or rocm (Windows)
- **NVIDIA GPUs**: cuda
- **NVIDIA GPUs**: cuda; vulkan available as explicit option
- **Intel GPUs**: vulkan (Linux/macOS) or sycl (Windows); openvino available as explicit option
- **Ascend NPUs**: cann
- **MUSA GPUs**: musa
Expand All @@ -69,7 +69,7 @@ Available backends depend on the detected GPU hardware.
- On **Windows**, vulkan is not supported on WSL2, so vendor-specific backends (rocm, sycl) are preferred

**Explicit backend selection**:
- **vulkan**: Use Vulkan-based inference (compatible with AMD, Intel, and CPU)
- **vulkan**: Use Vulkan-based inference (compatible with AMD, NVIDIA, Intel, and CPU)
- **rocm**: Use AMD ROCm backend (AMD GPUs only)
- **cuda**: Use NVIDIA CUDA backend (NVIDIA GPUs only)
- **sycl**: Use Intel SYCL/oneAPI backend (Intel GPUs only)
Expand Down
4 changes: 2 additions & 2 deletions docs/ramalama-serve.1.md
Original file line number Diff line number Diff line change
Expand Up @@ -87,7 +87,7 @@ Available backends depend on the detected GPU hardware.

**auto** (default): Automatically selects the preferred backend based on your GPU:
- **AMD GPUs**: vulkan (Linux/macOS) or rocm (Windows)
- **NVIDIA GPUs**: cuda
- **NVIDIA GPUs**: cuda; vulkan available as explicit option
- **Intel GPUs**: vulkan (Linux/macOS) or sycl (Windows); openvino available as explicit option
- **Ascend NPUs**: cann
- **MUSA GPUs**: musa
Expand All @@ -98,7 +98,7 @@ Available backends depend on the detected GPU hardware.
- On **Windows**, vulkan is not supported on WSL2, so vendor-specific backends (rocm, sycl) are preferred

**Explicit backend selection**:
- **vulkan**: Use Vulkan-based inference (compatible with AMD, Intel, and CPU)
- **vulkan**: Use Vulkan-based inference (compatible with AMD, NVIDIA, Intel, and CPU)
- **rocm**: Use AMD ROCm backend (AMD GPUs only)
- **cuda**: Use NVIDIA CUDA backend (NVIDIA GPUs only)
- **sycl**: Use Intel SYCL/oneAPI backend (Intel GPUs only)
Expand Down
4 changes: 2 additions & 2 deletions docs/ramalama.conf
Original file line number Diff line number Diff line change
Expand Up @@ -241,12 +241,12 @@
# Valid options: auto, vulkan, rocm, cuda, sycl, openvino, cann, musa
# - auto (default): Automatically selects the preferred backend based on detected GPU
# - AMD GPUs: vulkan (Linux/macOS) or rocm (Windows)
# - NVIDIA GPUs: cuda
# - NVIDIA GPUs: cuda; vulkan available as explicit option
# - Intel GPUs: vulkan (Linux/macOS) or sycl (Windows); openvino available as explicit option
# - Ascend NPUs: cann
# - MUSA GPUs: musa
# - No GPU: vulkan (CPU fallback)
# - vulkan: Use Vulkan-based inference (compatible with AMD, Intel, and CPU)
# - vulkan: Use Vulkan-based inference (compatible with AMD, NVIDIA, Intel, and CPU)
# - rocm: Use AMD ROCm backend (AMD GPUs only)
# - cuda: Use NVIDIA CUDA backend (NVIDIA GPUs only)
# - sycl: Use Intel SYCL/oneAPI backend (Intel GPUs only)
Expand Down
4 changes: 2 additions & 2 deletions docs/ramalama.conf.5.md
Original file line number Diff line number Diff line change
Expand Up @@ -240,13 +240,13 @@ Valid options: `auto`, `vulkan`, `rocm`, `cuda`, `sycl`, `openvino`, `cann`, `mu

- **auto** (default): Automatically selects the preferred backend based on detected GPU:
- AMD GPUs: vulkan (Linux/macOS) or rocm (Windows)
- NVIDIA GPUs: cuda
- NVIDIA GPUs: cuda; vulkan available as explicit option
- Intel GPUs: vulkan (Linux/macOS) or sycl (Windows); openvino available as explicit option
- Ascend NPUs: cann
- MUSA GPUs: musa
- No GPU: vulkan (CPU fallback)

- **vulkan**: Use Vulkan-based inference (compatible with AMD, Intel, and CPU)
- **vulkan**: Use Vulkan-based inference (compatible with AMD, NVIDIA, Intel, and CPU)
- **rocm**: Use AMD ROCm backend (AMD GPUs only)
- **cuda**: Use NVIDIA CUDA backend (NVIDIA GPUs only)
- **sycl**: Use Intel SYCL/oneAPI backend (Intel GPUs only)
Expand Down
27 changes: 27 additions & 0 deletions ramalama/cli.py
Original file line number Diff line number Diff line change
Expand Up @@ -125,6 +125,15 @@ def parse_port_option(option: str) -> str:


class OverrideDefaultAction(argparse.Action):
# Options whose default is computed while the parser is built, before the
# rest of the command line has been seen. Recorded so such a default can be
# told apart from a value the user passed.
dests: set[str] = set()

def __init__(self, *args, **kwargs):
Comment thread
olliewalsh marked this conversation as resolved.
super().__init__(*args, **kwargs)
OverrideDefaultAction.dests.add(self.dest)

def __call__(self, parser, namespace, values, option_string=None):
setattr(namespace, self.dest, values)
setattr(namespace, self.dest + '_override', True)
Expand Down Expand Up @@ -221,11 +230,29 @@ def parse_args_from_cmd(cmd: list[str]) -> tuple[argparse.ArgumentParser, argpar
post_parse_setup(args)

for arg in args.__dict__.keys() & config._fields:
if arg in OverrideDefaultAction.dests:
# These options record whether they were passed, so go by that rather
# than by the value. A computed default is not the user's choice and
# does not belong in the config: --image defaults to the image the
# detected GPU resolves to, worked out before --backend was parsed,
# and storing it marks the image as chosen rather than derived. A
# value the user did pass belongs there even where it matches the
# default, which is how "--image $RAMALAMA_DEFAULT_IMAGE" asks for
# that image instead of the one the GPU would select.
if getattr(args, f"{arg}_override", False):
setattr(config, arg, getattr(args, arg))
continue
if getattr(args, arg) != getattr(config, arg):
setattr(config, arg, getattr(args, arg))

runtime_plugin = get_runtime(config.runtime)
runtime_plugin.sync_args_to_runtime_config(args, config)

# The runtime config now carries --backend, so the image the command will
# run can be resolved for real.
if hasattr(args, "image") and not getattr(args, "image_override", False):
args.image = accel_image(config)

return parser, args


Expand Down
50 changes: 47 additions & 3 deletions ramalama/common.py
Original file line number Diff line number Diff line change
Expand Up @@ -57,6 +57,22 @@ def sanitize_filename(filename: str) -> str:

podman_machine_accel = False

# CDI device names of the NVIDIA GPUs the user narrowed CUDA_VISIBLE_DEVICES
# down to, set by check_nvidia(). Empty when every detected GPU is in play.
nvidia_selected_devices: list[str] = []


def container_cuda_visible_devices(value: str) -> str:
"""CUDA_VISIBLE_DEVICES as it should read inside the container.

Where the selection is narrowed, only those GPUs are passed in and the
container numbers them from zero, so the host's indices no longer describe
them. Returns the value unchanged when every GPU is in play.
"""
if not nvidia_selected_devices:
return value
return ",".join(str(i) for i in range(len(nvidia_selected_devices)))


def confirm_no_gpu(name, provider) -> bool:
while True:
Expand Down Expand Up @@ -438,9 +454,13 @@ def find_in_cdi(devices: list[str]) -> tuple[list[str], list[str]]:
for device in devices:
if device in cdi_device_names:
configured.append(device)
# A device can be specified by a prefix of the uuid
elif device.startswith("GPU") and any(name.startswith(device) for name in cdi_device_names):
configured.append(device)
# A device can be specified by a prefix of the uuid. Record the full
# name it resolves to: it reaches "--device nvidia.com/gpu=<name>",
# which matches the CDI configuration exactly and not by prefix.
elif device.startswith("GPU") and (
full_name := next((name for name in cdi_device_names if name.startswith(device)), None)
):
configured.append(full_name)
else:
perror(f"Device {device} does not have a CDI configuration")
unconfigured.append(device)
Expand Down Expand Up @@ -468,6 +488,20 @@ def check_metal(args: ContainerArgType) -> bool:
return platform.system() == "Darwin"


@lru_cache(maxsize=1)
def has_nvidia_vulkan_icd() -> bool:
"""True when NVIDIA's Vulkan ICD manifest is installed on the host.

The container toolkit injects the ICD from the host driver installation, so
when it is missing here the vulkan backend finds only mesa's llvmpipe inside
the container and runs on the CPU instead of failing.
"""
return any(
glob.glob(os.path.join(host_path(icd_dir), "*nvidia*.json"))
for icd_dir in ("/usr/share/vulkan/icd.d", "/etc/vulkan/icd.d")
)


@lru_cache(maxsize=1)
def check_nvidia() -> Optional[Literal["cuda"]]:
try:
Expand Down Expand Up @@ -515,6 +549,16 @@ def check_nvidia() -> Optional[Literal["cuda"]]:
if not configured:
configured = indices

# Record a narrowed selection so the engine can pass just those GPUs
# into the container. Not every backend filters on CUDA_VISIBLE_DEVICES
# (llama.cpp's Vulkan backend indexes Vulkan's own device enumeration),
# so the selection has to happen at the device level to be honoured.
# These names came back from find_in_cdi(), so they are known to the CDI
# configuration; the fallback above to every index has not been checked
# and stays on the "all" device.
global nvidia_selected_devices
nvidia_selected_devices = configured if set(configured) != set(indices) else []
Comment thread
olliewalsh marked this conversation as resolved.

os.environ["CUDA_VISIBLE_DEVICES"] = ','.join(configured)
return "cuda"

Expand Down
28 changes: 23 additions & 5 deletions ramalama/compose.py
Original file line number Diff line number Diff line change
Expand Up @@ -5,7 +5,9 @@
import shlex
from typing import Optional

from ramalama.common import RAG_DIR, get_accel_env_vars, get_gpu_devices
# Live reference for checking global vars
import ramalama.common
from ramalama.common import RAG_DIR, container_cuda_visible_devices, get_accel_env_vars, get_gpu_devices
from ramalama.file import PlainFile
from ramalama.host_utils import format_bind_host_publish_prefix
from ramalama.version import version
Expand Down Expand Up @@ -122,6 +124,10 @@ def _gen_ports(self) -> str:

def _gen_environment(self) -> str:
env_vars = get_accel_env_vars()
if "CUDA_VISIBLE_DEVICES" in env_vars:
# device_ids below reserves just the selected GPUs, so the container
# renumbers them, exactly as it does for "ramalama run".
env_vars["CUDA_VISIBLE_DEVICES"] = container_cuda_visible_devices(env_vars["CUDA_VISIBLE_DEVICES"])
# Allow user to override with --env
if getattr(self.args, "env", None):
for e in self.args.env:
Expand All @@ -137,17 +143,29 @@ def _gen_environment(self) -> str:
return env_spec

def _gen_gpu_deployment(self) -> str:
gpu_keywords = ["cuda", "rocm", "gpu"]
if not any(keyword in self.image.lower() for keyword in gpu_keywords):
# The "nvidia" device driver only covers NVIDIA GPUs, so key the
# reservation off the detected hardware. The image name does not
# identify it: NVIDIA can be served by the vulkan-capable ramalama
# image as well as by the cuda one.
if "CUDA_VISIBLE_DEVICES" not in get_accel_env_vars():
return ""

return """\
# Reserve only the GPUs the user selected. CUDA_VISIBLE_DEVICES alone
# cannot narrow it: the Vulkan backend never reads that variable, so the
# other GPUs have to be kept out of the container entirely.
if selected := ramalama.common.nvidia_selected_devices:
device_ids = ", ".join(f'"{name}"' for name in selected)
reservation = f"device_ids: [{device_ids}]"
else:
reservation = "count: all"

return f"""\
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: all
{reservation}
capabilities: [gpu]"""

def _gen_command(self) -> str:
Expand Down
Loading
Loading