Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
9 changes: 6 additions & 3 deletions docs/options/backend.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,16 +9,19 @@ GPU backend to use for inference (default: auto).
Available backends depend on the detected GPU hardware.

**auto** (default): Automatically selects the preferred backend based on your GPU:
- **AMD GPUs**: vulkan (Linux/macOS) or rocm (Windows)
- **AMD GPUs**: vulkan (Linux/macOS) or rocm (WSL2)
- **NVIDIA GPUs**: cuda; vulkan available as explicit option
- **Intel GPUs**: vulkan (Linux/macOS) or sycl (Windows); openvino available as explicit option
- **Intel GPUs**: vulkan (Linux/macOS) or sycl (WSL2); openvino available as explicit option
- **Ascend NPUs**: cann
- **MUSA GPUs**: musa
- **No GPU**: vulkan (CPU fallback)

**Platform-specific behavior**:
- On **Linux/macOS**, Vulkan provides broad compatibility and is preferred for AMD and Intel GPUs
- On **Windows**, vulkan is not supported on WSL2, so vendor-specific backends (rocm, sycl) are preferred
- On **WSL2**, Vulkan is a poor default, so vendor-specific backends (rocm, sycl) are preferred. It
remains available, and `--backend=vulkan` still selects it. This covers both Windows, where
containers run in the WSL2-backed machine, and ramalama running inside a WSL2 distro, which
otherwise looks like Linux

**Explicit backend selection**:
- **vulkan**: Use Vulkan-based inference (compatible with AMD, NVIDIA, Intel, and CPU)
Expand Down
9 changes: 6 additions & 3 deletions docs/ramalama-bench.1.md
Original file line number Diff line number Diff line change
Expand Up @@ -45,16 +45,19 @@ GPU backend to use for inference (default: auto).
Available backends depend on the detected GPU hardware.

**auto** (default): Automatically selects the preferred backend based on your GPU:
- **AMD GPUs**: vulkan (Linux/macOS) or rocm (Windows)
- **AMD GPUs**: vulkan (Linux/macOS) or rocm (WSL2)
- **NVIDIA GPUs**: cuda; vulkan available as explicit option
- **Intel GPUs**: vulkan (Linux/macOS) or sycl (Windows); openvino available as explicit option
- **Intel GPUs**: vulkan (Linux/macOS) or sycl (WSL2); openvino available as explicit option
- **Ascend NPUs**: cann
- **MUSA GPUs**: musa
- **No GPU**: vulkan (CPU fallback)

**Platform-specific behavior**:
- On **Linux/macOS**, Vulkan provides broad compatibility and is preferred for AMD and Intel GPUs
- On **Windows**, vulkan is not supported on WSL2, so vendor-specific backends (rocm, sycl) are preferred
- On **WSL2**, Vulkan is a poor default, so vendor-specific backends (rocm, sycl) are preferred. It
remains available, and `--backend=vulkan` still selects it. This covers both Windows, where
containers run in the WSL2-backed machine, and ramalama running inside a WSL2 distro, which
otherwise looks like Linux

**Explicit backend selection**:
- **vulkan**: Use Vulkan-based inference (compatible with AMD, NVIDIA, Intel, and CPU)
Expand Down
2 changes: 2 additions & 0 deletions docs/ramalama-cuda.7.md
Original file line number Diff line number Diff line change
Expand Up @@ -138,6 +138,8 @@ ramalama run granite

This is particularly useful in multi-GPU systems where you want to dedicate specific GPUs to different workloads.

Where the CDI configuration has an entry for each GPU, only the selected ones are passed into the container by the NVIDIA container toolkit; where it only defines `all`, every detected GPU is. Either way the host's other GPU devices, such as `/dev/dri` for an integrated GPU, are left out, since the Vulkan backend offloads onto every device it can enumerate and would otherwise use them. Pass `--device /dev/dri` to `ramalama run` or `ramalama serve` to add them back.

If `CUDA_VISIBLE_DEVICES` is set to an empty string, RamaLama treats it as unset and follows the default GPU-selection behavior.

```bash
Expand Down
9 changes: 6 additions & 3 deletions docs/ramalama-perplexity.1.md
Original file line number Diff line number Diff line change
Expand Up @@ -45,16 +45,19 @@ GPU backend to use for inference (default: auto).
Available backends depend on the detected GPU hardware.

**auto** (default): Automatically selects the preferred backend based on your GPU:
- **AMD GPUs**: vulkan (Linux/macOS) or rocm (Windows)
- **AMD GPUs**: vulkan (Linux/macOS) or rocm (WSL2)
- **NVIDIA GPUs**: cuda; vulkan available as explicit option
- **Intel GPUs**: vulkan (Linux/macOS) or sycl (Windows); openvino available as explicit option
- **Intel GPUs**: vulkan (Linux/macOS) or sycl (WSL2); openvino available as explicit option
- **Ascend NPUs**: cann
- **MUSA GPUs**: musa
- **No GPU**: vulkan (CPU fallback)

**Platform-specific behavior**:
- On **Linux/macOS**, Vulkan provides broad compatibility and is preferred for AMD and Intel GPUs
- On **Windows**, vulkan is not supported on WSL2, so vendor-specific backends (rocm, sycl) are preferred
- On **WSL2**, Vulkan is a poor default, so vendor-specific backends (rocm, sycl) are preferred. It
remains available, and `--backend=vulkan` still selects it. This covers both Windows, where
containers run in the WSL2-backed machine, and ramalama running inside a WSL2 distro, which
otherwise looks like Linux

**Explicit backend selection**:
- **vulkan**: Use Vulkan-based inference (compatible with AMD, NVIDIA, Intel, and CPU)
Expand Down
9 changes: 6 additions & 3 deletions docs/ramalama-run.1.md
Original file line number Diff line number Diff line change
Expand Up @@ -57,16 +57,19 @@ GPU backend to use for inference (default: auto).
Available backends depend on the detected GPU hardware.

**auto** (default): Automatically selects the preferred backend based on your GPU:
- **AMD GPUs**: vulkan (Linux/macOS) or rocm (Windows)
- **AMD GPUs**: vulkan (Linux/macOS) or rocm (WSL2)
- **NVIDIA GPUs**: cuda; vulkan available as explicit option
- **Intel GPUs**: vulkan (Linux/macOS) or sycl (Windows); openvino available as explicit option
- **Intel GPUs**: vulkan (Linux/macOS) or sycl (WSL2); openvino available as explicit option
- **Ascend NPUs**: cann
- **MUSA GPUs**: musa
- **No GPU**: vulkan (CPU fallback)

**Platform-specific behavior**:
- On **Linux/macOS**, Vulkan provides broad compatibility and is preferred for AMD and Intel GPUs
- On **Windows**, vulkan is not supported on WSL2, so vendor-specific backends (rocm, sycl) are preferred
- On **WSL2**, Vulkan is a poor default, so vendor-specific backends (rocm, sycl) are preferred. It
remains available, and `--backend=vulkan` still selects it. This covers both Windows, where
containers run in the WSL2-backed machine, and ramalama running inside a WSL2 distro, which
otherwise looks like Linux

**Explicit backend selection**:
- **vulkan**: Use Vulkan-based inference (compatible with AMD, NVIDIA, Intel, and CPU)
Expand Down
9 changes: 6 additions & 3 deletions docs/ramalama-sandbox-goose.1.md
Original file line number Diff line number Diff line change
Expand Up @@ -53,16 +53,19 @@ GPU backend to use for inference (default: auto).
Available backends depend on the detected GPU hardware.

**auto** (default): Automatically selects the preferred backend based on your GPU:
- **AMD GPUs**: vulkan (Linux/macOS) or rocm (Windows)
- **AMD GPUs**: vulkan (Linux/macOS) or rocm (WSL2)
- **NVIDIA GPUs**: cuda; vulkan available as explicit option
- **Intel GPUs**: vulkan (Linux/macOS) or sycl (Windows); openvino available as explicit option
- **Intel GPUs**: vulkan (Linux/macOS) or sycl (WSL2); openvino available as explicit option
- **Ascend NPUs**: cann
- **MUSA GPUs**: musa
- **No GPU**: vulkan (CPU fallback)

**Platform-specific behavior**:
- On **Linux/macOS**, Vulkan provides broad compatibility and is preferred for AMD and Intel GPUs
- On **Windows**, vulkan is not supported on WSL2, so vendor-specific backends (rocm, sycl) are preferred
- On **WSL2**, Vulkan is a poor default, so vendor-specific backends (rocm, sycl) are preferred. It
remains available, and `--backend=vulkan` still selects it. This covers both Windows, where
containers run in the WSL2-backed machine, and ramalama running inside a WSL2 distro, which
otherwise looks like Linux

**Explicit backend selection**:
- **vulkan**: Use Vulkan-based inference (compatible with AMD, NVIDIA, Intel, and CPU)
Expand Down
9 changes: 6 additions & 3 deletions docs/ramalama-sandbox-opencode.1.md
Original file line number Diff line number Diff line change
Expand Up @@ -53,16 +53,19 @@ GPU backend to use for inference (default: auto).
Available backends depend on the detected GPU hardware.

**auto** (default): Automatically selects the preferred backend based on your GPU:
- **AMD GPUs**: vulkan (Linux/macOS) or rocm (Windows)
- **AMD GPUs**: vulkan (Linux/macOS) or rocm (WSL2)
- **NVIDIA GPUs**: cuda; vulkan available as explicit option
- **Intel GPUs**: vulkan (Linux/macOS) or sycl (Windows); openvino available as explicit option
- **Intel GPUs**: vulkan (Linux/macOS) or sycl (WSL2); openvino available as explicit option
- **Ascend NPUs**: cann
- **MUSA GPUs**: musa
- **No GPU**: vulkan (CPU fallback)

**Platform-specific behavior**:
- On **Linux/macOS**, Vulkan provides broad compatibility and is preferred for AMD and Intel GPUs
- On **Windows**, vulkan is not supported on WSL2, so vendor-specific backends (rocm, sycl) are preferred
- On **WSL2**, Vulkan is a poor default, so vendor-specific backends (rocm, sycl) are preferred. It
remains available, and `--backend=vulkan` still selects it. This covers both Windows, where
containers run in the WSL2-backed machine, and ramalama running inside a WSL2 distro, which
otherwise looks like Linux

**Explicit backend selection**:
- **vulkan**: Use Vulkan-based inference (compatible with AMD, NVIDIA, Intel, and CPU)
Expand Down
9 changes: 6 additions & 3 deletions docs/ramalama-sandbox-pi.1.md
Original file line number Diff line number Diff line change
Expand Up @@ -57,16 +57,19 @@ GPU backend to use for inference (default: auto).
Available backends depend on the detected GPU hardware.

**auto** (default): Automatically selects the preferred backend based on your GPU:
- **AMD GPUs**: vulkan (Linux/macOS) or rocm (Windows)
- **AMD GPUs**: vulkan (Linux/macOS) or rocm (WSL2)
- **NVIDIA GPUs**: cuda; vulkan available as explicit option
- **Intel GPUs**: vulkan (Linux/macOS) or sycl (Windows); openvino available as explicit option
- **Intel GPUs**: vulkan (Linux/macOS) or sycl (WSL2); openvino available as explicit option
- **Ascend NPUs**: cann
- **MUSA GPUs**: musa
- **No GPU**: vulkan (CPU fallback)

**Platform-specific behavior**:
- On **Linux/macOS**, Vulkan provides broad compatibility and is preferred for AMD and Intel GPUs
- On **Windows**, vulkan is not supported on WSL2, so vendor-specific backends (rocm, sycl) are preferred
- On **WSL2**, Vulkan is a poor default, so vendor-specific backends (rocm, sycl) are preferred. It
remains available, and `--backend=vulkan` still selects it. This covers both Windows, where
containers run in the WSL2-backed machine, and ramalama running inside a WSL2 distro, which
otherwise looks like Linux

**Explicit backend selection**:
- **vulkan**: Use Vulkan-based inference (compatible with AMD, NVIDIA, Intel, and CPU)
Expand Down
9 changes: 6 additions & 3 deletions docs/ramalama-serve.1.md
Original file line number Diff line number Diff line change
Expand Up @@ -86,16 +86,19 @@ GPU backend to use for inference (default: auto).
Available backends depend on the detected GPU hardware.

**auto** (default): Automatically selects the preferred backend based on your GPU:
- **AMD GPUs**: vulkan (Linux/macOS) or rocm (Windows)
- **AMD GPUs**: vulkan (Linux/macOS) or rocm (WSL2)
- **NVIDIA GPUs**: cuda; vulkan available as explicit option
- **Intel GPUs**: vulkan (Linux/macOS) or sycl (Windows); openvino available as explicit option
- **Intel GPUs**: vulkan (Linux/macOS) or sycl (WSL2); openvino available as explicit option
- **Ascend NPUs**: cann
- **MUSA GPUs**: musa
- **No GPU**: vulkan (CPU fallback)

**Platform-specific behavior**:
- On **Linux/macOS**, Vulkan provides broad compatibility and is preferred for AMD and Intel GPUs
- On **Windows**, vulkan is not supported on WSL2, so vendor-specific backends (rocm, sycl) are preferred
- On **WSL2**, Vulkan is a poor default, so vendor-specific backends (rocm, sycl) are preferred. It
remains available, and `--backend=vulkan` still selects it. This covers both Windows, where
containers run in the WSL2-backed machine, and ramalama running inside a WSL2 distro, which
otherwise looks like Linux

**Explicit backend selection**:
- **vulkan**: Use Vulkan-based inference (compatible with AMD, NVIDIA, Intel, and CPU)
Expand Down
11 changes: 6 additions & 5 deletions docs/ramalama.conf
Original file line number Diff line number Diff line change
Expand Up @@ -240,9 +240,9 @@
#
# Valid options: auto, vulkan, rocm, cuda, sycl, openvino, cann, musa
# - auto (default): Automatically selects the preferred backend based on detected GPU
# - AMD GPUs: vulkan (Linux/macOS) or rocm (Windows)
# - AMD GPUs: vulkan (Linux/macOS) or rocm (WSL2)
# - NVIDIA GPUs: cuda; vulkan available as explicit option
# - Intel GPUs: vulkan (Linux/macOS) or sycl (Windows); openvino available as explicit option
# - Intel GPUs: vulkan (Linux/macOS) or sycl (WSL2); openvino available as explicit option
# - Ascend NPUs: cann
# - MUSA GPUs: musa
# - No GPU: vulkan (CPU fallback)
Expand All @@ -254,9 +254,10 @@
# - cann: Use Huawei CANN backend (Ascend NPUs only); uses quay.io/ramalama/cann
# - musa: Use Moore Threads MUSA backend (MUSA GPUs only); uses quay.io/ramalama/musa
#
# Platform-specific behavior: On Windows, vulkan is not supported on WSL2, so
# vendor-specific backends (rocm for AMD, sycl for Intel) are automatically
# preferred when using backend="auto".
# Platform-specific behavior: vulkan is not supported on WSL2, so vendor-specific
# backends (rocm for AMD, sycl for Intel) are automatically preferred when using
# backend="auto". This covers both Windows and ramalama running inside a WSL2
# distro, which otherwise looks like Linux.
#
#backend = "auto"

Expand Down
6 changes: 3 additions & 3 deletions docs/ramalama.conf.5.md
Original file line number Diff line number Diff line change
Expand Up @@ -239,9 +239,9 @@ This setting affects which container image is selected and how GPU resources are
Valid options: `auto`, `vulkan`, `rocm`, `cuda`, `sycl`, `openvino`, `cann`, `musa`.

- **auto** (default): Automatically selects the preferred backend based on detected GPU:
- AMD GPUs: vulkan (Linux/macOS) or rocm (Windows)
- AMD GPUs: vulkan (Linux/macOS) or rocm (WSL2)
- NVIDIA GPUs: cuda; vulkan available as explicit option
- Intel GPUs: vulkan (Linux/macOS) or sycl (Windows); openvino available as explicit option
- Intel GPUs: vulkan (Linux/macOS) or sycl (WSL2); openvino available as explicit option
- Ascend NPUs: cann
- MUSA GPUs: musa
- No GPU: vulkan (CPU fallback)
Expand All @@ -254,7 +254,7 @@ Valid options: `auto`, `vulkan`, `rocm`, `cuda`, `sycl`, `openvino`, `cann`, `mu
- **cann**: Use Huawei CANN backend (Ascend NPUs only); uses `quay.io/ramalama/cann`
- **musa**: Use Moore Threads MUSA backend (MUSA GPUs only); uses `quay.io/ramalama/musa`

**Platform-specific behavior**: On Windows, vulkan is not supported on WSL2, so vendor-specific backends (rocm for AMD, sycl for Intel) are automatically preferred when using `backend="auto"`.
**Platform-specific behavior**: vulkan is not supported on WSL2, so vendor-specific backends (rocm for AMD, sycl for Intel) are automatically preferred when using `backend="auto"`. This covers both Windows and ramalama running inside a WSL2 distro, which otherwise looks like Linux.

Example configuration:

Expand Down
43 changes: 41 additions & 2 deletions ramalama/common.py
Original file line number Diff line number Diff line change
Expand Up @@ -11,7 +11,7 @@
import string
import subprocess
import sys
from collections.abc import Callable, Sequence
from collections.abc import Callable, Collection, Sequence
from dataclasses import dataclass
from functools import lru_cache
from pathlib import Path
Expand Down Expand Up @@ -488,6 +488,33 @@ def check_metal(args: ContainerArgType) -> bool:
return platform.system() == "Darwin"


@lru_cache(maxsize=1)
def in_wsl() -> bool:
"""True when the interpreter itself is running inside a WSL distro.

False on native Windows, where ramalama drives a podman machine instead.
"""
try:
with open("/proc/sys/kernel/osrelease") as f:
return "microsoft" in f.read().lower()
except OSError:
return False


def is_windows_or_wsl() -> bool:
"""True where containers reach GPUs through WSL rather than native devices.

Covers both a native Windows interpreter, which runs containers in the
WSL2-backed podman machine, and ramalama running inside a WSL distro
itself, which platform.system() reports as "Linux".

WSL exposes GPUs through /dev/dxg, so Vulkan there means mesa's dzn driver
translating to D3D12 (no cooperative matrix support, no compute tuning) or
a silent llvmpipe fallback.
"""
return platform.system() == "Windows" or in_wsl()


@lru_cache(maxsize=1)
def has_nvidia_vulkan_icd() -> bool:
"""True when NVIDIA's Vulkan ICD manifest is installed on the host.
Expand Down Expand Up @@ -722,7 +749,19 @@ def set_gpu_type_env_vars():
]


def get_gpu_devices():
def get_gpu_devices(accel_env_vars: Optional[Collection[str]] = None) -> dict[str, str]:
"""The host GPU devices to hand to the container, given the accelerator in play.

An NVIDIA GPU does not come in this way: the container toolkit injects the
device nodes of the GPUs that were asked for, DRM nodes included. Mapping
the host's GPU devices in as well can then only add ones that are not the
accelerator in play - an iGPU on a hybrid host, or a GPU left out of a
narrowed selection - and llama.cpp's Vulkan backend offloads onto every
device it can enumerate. "--device" remains for anyone who wants them.
"""
if "CUDA_VISIBLE_DEVICES" in (get_gpu_type_env_vars() if accel_env_vars is None else accel_env_vars):
return {}

devices = {}
for dev in ["dri", "kfd", "accel"]:
path = "/dev/" + dev
Expand Down
2 changes: 1 addition & 1 deletion ramalama/compose.py
Original file line number Diff line number Diff line change
Expand Up @@ -98,7 +98,7 @@ def _gen_mmproj_volume(self) -> str:
return f'\n - "{self.src_mmproj_path}:{self.dest_mmproj_path}:ro"'

def _gen_devices(self) -> str:
devices = get_gpu_devices()
devices = get_gpu_devices(get_accel_env_vars())

if not devices:
return ""
Expand Down
23 changes: 16 additions & 7 deletions ramalama/engine.py
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,6 @@
import glob
import json
import os
import platform
import subprocess
import sys
import time
Expand All @@ -22,7 +21,9 @@
exec_cmd,
genname,
get_accel_env_vars,
get_gpu_devices,
host_path,
is_windows_or_wsl,
perror,
run_cmd,
)
Expand Down Expand Up @@ -120,12 +121,14 @@ def add_device_options(self):
if ramalama.common.podman_machine_accel:
self.exec_args += ["--device", "/dev/dri"]

for path in ["/dev/dri", "/dev/kfd", "/dev/accel", "/dev/davinci*", "/dev/devmm_svm", "/dev/hisi_hdc"]:
env_vars = get_accel_env_vars()
gpu_devices = list(get_gpu_devices(env_vars).values())
for path in [*gpu_devices, "/dev/davinci*", "/dev/devmm_svm", "/dev/hisi_hdc"]:
for dev in glob.glob(path):
self.exec_args += ["--device", dev]

intel_windows_added = False
for k, v in get_accel_env_vars().items():
wsl_devices_added = False
for k, v in env_vars.items():
# Special case for Cuda
if k == "CUDA_VISIBLE_DEVICES":
# Pass in only the GPUs the user selected rather than all of
Expand All @@ -147,11 +150,17 @@ def add_device_options(self):
v = container_cuda_visible_devices(v)
elif k == "MUSA_VISIBLE_DEVICES":
self.exec_args += ["--env", "MTHREADS_VISIBLE_DEVICES=all"]
elif k == "INTEL_VISIBLE_DEVICES":
if platform.system() == "Windows" and not intel_windows_added:
elif k in ("HIP_VISIBLE_DEVICES", "INTEL_VISIBLE_DEVICES"):
# WSL exposes the GPU as /dev/dxg with its driver libraries in
# /usr/lib/wsl, whether ramalama runs on native Windows against
# the podman machine or inside the distro itself, where
# platform.system() reports "Linux". That is how both the AMD
# and the Intel GPU come in, so the rocm backend needs it as
# much as sycl does.
if is_windows_or_wsl() and not wsl_devices_added:
self.exec_args += ["--device", "/dev/dxg"]
self.exec_args += ["--mount", "type=bind,src=/usr/lib/wsl,dst=/usr/lib/wsl"]
intel_windows_added = True
wsl_devices_added = True

self.exec_args += ["-e", f"{k}={v}"]

Expand Down
4 changes: 2 additions & 2 deletions ramalama/kube.py
Original file line number Diff line number Diff line change
Expand Up @@ -85,7 +85,7 @@ def _gen_volumes(self) -> Tuple[str, str]:
def _gen_devices(self) -> Tuple[str, str]:
mounts = ""
volumes = ""
for name, path in get_gpu_devices().items():
for name, path in get_gpu_devices(get_accel_env_vars()).items():
mounts += f"""
- mountPath: {path}
name: {name}"""
Expand Down Expand Up @@ -219,7 +219,7 @@ def __gen_resources(self) -> str:
limits:
'nvidia.com/gpu=all': 1"""

devices = get_gpu_devices()
devices = get_gpu_devices(get_accel_env_vars())
if devices:
limits = "".join(f"\n 'podman.io/device={path}': 1" for path in devices.values())
return f"""
Expand Down
Loading
Loading