Short notes for getting vla.cpp to compile and link on macOS. The Metal
backend is auto-detected by the vendored llama.cpp/ggml and needs no special
CMake flag.
brew install protobuf zeromq cppzmq pkg-configAll four are required at configure time, not optional: find_package(Protobuf)
and pkg_check_modules(libzmq) are unconditional, so a missing zeromq or
cppzmq fails CMake before anything builds.
To let the binaries fetch checkpoints with -hf, also install the Hugging Face
CLI - the fetch shells out to hf and stops with hf: command not found
without it:
pip install -U "huggingface_hub[cli]" # or: uv tool install huggingface_hubOn MacOS, Metal is enabled by default. Using Metal makes the computation run on the GPU.
To disable the Metal build at compile time use the -DGGML_METAL=OFF cmake option.
# cmake fetches llama.cpp at the pinned tag; no patch step.
# On MacOS, Metal is enabled by default
cmake -B build -DCMAKE_BUILD_TYPE=Release
cmake --build build -j$(sysctl -n hw.ncpu)The VLA binaries pick their backend at load time and take no flag for it:
-DGGML_METAL=OFF at configure time is the only way to put vla-cli,
vla-server and vla-bench on the CPU. VLA_DEVICE only chooses an ordinal
for CUDA and SYCL, so it does nothing here, and --n-gpu-layers belongs to
vlm-server on the VLM path, not to the VLA binaries.
Each arch calls backend_init (src/backend.h) exactly once at load time and
runs everything on what it returns, vision tower included. On macOS that is
Metal (ggml_backend_metal_init). Confirm it from the startup banner, which is
tagged with the arch:
vla(pi0): backend = Metal
SmolVLA is the one that logs under a bare vla tag, so for it the line reads:
vla: backend = Metal
There is no second banner to look for: the VLA path prints no clip_ctx: CLIP using GPU backend line, because there is no separate CLIP context to bring up -
the single backend already covers both towers. A Metal build that reports only
the line above is working.
If you instead see vla(<arch>): backend = CPU (N threads), the build didn't
pick up Metal - rebuild from a clean build/ and check GGML_METAL is ON in
the CMake cache (grep GGML_METAL build/CMakeCache.txt).
Single-backend, no per-op CPU fallback: the core uses one backend +
gallocr, not a scheduler. SmolVLA's ops are all Metal-supported; an arch that hits an unimplemented op would assert at predict time rather than silently fall back.
BitVLA is the exception and does not run on Metal at all. It calls
ggml_backend_cpu_init() directly (src/models/bitvla.cpp:568) because its
graph stays on CPU and the LM offloads through CUDA, so it reports vla(bitvla): ggml backend = CPU (N threads) even on a Metal build - that banner is expected,
not a broken build. The published GGUFs are also int2-packed, which model_load
rejects outside a CUDA build (VLA_BITVLA_CUDA_KERNELS), so on macOS it fails
to load rather than running slowly.
vla-bench times predict() in-process on synthetic inputs: engine only, no
transport, no simulator, no claim about task success. Apple M5 Max (18-core CPU,
40-core GPU, 64 GB unified memory), macOS 26.6.1, High Power Mode, llama.cpp
b10331, weights as shipped, 20 reps after 3 warmups, best of three sweeps (the
sweep with the lowest p50), each model at its native input size and view count -
the same counts the
RTX 5090 report uses, so the two compare cell for cell.
| Model | Views | Input | min ms | p50 ms | p90 ms | vision ms |
|---|---|---|---|---|---|---|
| VLA-Adapter | 1 | 224 | 64.5 | 64.8 | 65.1 | 33.0 |
| VLA-JEPA | 1 | 256 | 74.9 | 75.1 | 75.4 | 17.2 |
| GR00T N1.5 | 1 | 224 | 104.0 | 104.3 | 104.6 | 20.7 |
| SmolVLA | 2 | 512 | 114.8 | 115.2 | 115.9 | 29.0 |
| GR00T N1.7 | 1 | 256 | 128.0 | 128.4 | 128.9 | 17.3 |
| GR00T N1.6 | 1 | 224 | 133.0 | 133.3 | 134.4 | 21.3 |
| OpenVLA-OFT | 1 | 224 | 183.4 | 184.2 | 184.6 | 33.7 |
| Evo-1 | 1 | 448 | 214.5 | 215.0 | 216.2 | 50.3 |
| pi0 | 2 | 224 | 220.2 | 220.8 | 221.3 | 39.3 |
| pi0.5 | 2 | 224 | 237.1 | 237.5 | 237.7 | 39.1 |
Every run came up on Metal (backend = Metal in the load banner); none fell
back to CPU. The three sweeps agree to within 0.6% per model, and min to p90
spans no more than 2 ms, so these settle rather than scatter.
BitVLA has no row, for the reason in the section above: there is no Metal path to time.
Outputs were checked against the CPU backend on all ten models above, on an M5
Max, using vla_predict_check on fixed images, tokens, state and noise - the
same build twice, once as configured and once with -DGGML_METAL=OFF. It is a
test target, so add -DVLA_BUILD_TESTS=ON to the configure line above to get
it.
The loosest model is π0: 1.9e-2 absolute on actions peaking near 1.0, RMS 5.2e-4 for that model, with a worst per-model RMS of 1.7e-3 across the set. Eight of the ten stay under 7e-3 absolute. The two outliers are π0 and GR00T N1.7 (1.4e-2), which run multi-step denoise loops where per-step rounding compounds. Metal output is deterministic run to run - repeated runs are bit-identical - so these are BF16/F32 kernel rounding, not instability.
These predate the table above and are not the same experiment - they were taken
end-to-end through vla-server on an M4, so they include transport and
preprocessing that vla-bench excludes, and they come from an older revision.
Read them as evidence that GPU offload is worth having, not as current numbers.
SmolVLA (libero, mmproj + 878 MiB BF16 weights), Apple M4, steady state:
| Stage | CPU (before) | Metal GPU (after) |
|---|---|---|
| vision | 22,367 ms | ~178 ms |
| inference | 12,878 ms | ~144 ms |
| total/req | ~35,250 ms | ~324 ms |
≈ 108× faster end-to-end. First request is ~671 ms (Metal pipeline warmup), then it settles to ~321–328 ms/req.
| Model | SR | Client/step | Server/step |
|---|---|---|---|
| SmolVLA | 0.7 | 888 ms | 324 ms (181 vision + 141 inf + 2) |
| Pi0 | 0.8 | 1135 ms | 1129 ms (922 vision + 200 inf + 7) |
| Gr00t-n1.7 | 1.0 | 755 ms | 600 ms (185 vision + 405 inf + 10) |