Skip to content

Commit 39d1eea

Browse files
committed
record seven supported archs, the gelu fix and the corrected gr00t n1.7 findings
1 parent 57e5f31 commit 39d1eea

2 files changed

Lines changed: 104 additions & 77 deletions

File tree

‎README.md‎

Lines changed: 3 additions & 10 deletions
Original file line numberDiff line numberDiff line change
@@ -312,26 +312,19 @@ Experimental results on other platforms can be found in
312312
Support matrix of models (rows) against platforms (columns). Legend: `Y` =
313313
supported (released and benchmarked), `~` = in progress, `-` = planned.
314314

315-
OpenVINO runs SmolVLA, π0.5, Evo-1, VLA-Adapter and GR00T N1.6 on Intel CPUs,
316-
GPUs and NPUs, matching the CPU backend to 3e-3; on an Arc B390 iGPU that is 3.1x
317-
to 8.2x the native CPU backend. VLA-JEPA and GR00T N1.5 run but drift further and
318-
are marked in progress; GR00T N1.7 is wrong and stays unsupported. Read the known issues in
319-
[docs/backend/ov.md](docs/backend/ov.md) before running it - in particular, do
320-
not set `GGML_OPENVINO_CACHE_DIR`.
321-
322-
| Model | CPU (x86-64 / ARM) | CUDA | SYCL (Intel) | Metal | OpenVINO |
315+
| Model | CPU (x86-64 / ARM) | CUDA | [SYCL (Intel)](docs/backend/sycl.md) | [Metal](docs/backend/metal.md) | [OpenVINO](docs/backend/ov.md) |
323316
|---|:--:|:--:|:--:|:--:|:--:|
324317
| [SmolVLA](https://hf.co/vrfai/smolvla-libero-gguf) | Y | Y | Y | Y | Y |
325318
| [π0](https://hf.co/vrfai/pi0-libero-finetuned-v044-gguf) | Y | Y | - | Y | ~ |
326319
| [π0.5](https://hf.co/vrfai/pi05-libero-gguf) | Y | Y | - | Y | Y |
327-
| [GR00T N1.5](https://hf.co/vrfai/gr00tn1d5-libero-object-gguf) | Y | Y | - | Y | ~ |
320+
| [GR00T N1.5](https://hf.co/vrfai/gr00tn1d5-libero-object-gguf) | Y | Y | - | Y | Y |
328321
| [GR00T N1.6](https://hf.co/vrfai/gr00tn1d6-libero-gguf) | Y | Y | - | Y | Y |
329322
| [GR00T N1.7](https://hf.co/vrfai/gr00tn1d7-libero-gguf) | Y | Y | - | Y | - |
330323
| [BitVLA](https://hf.co/vrfai/bitvla-libero-gguf) | Y | Y | - | ~ | - |
331324
| [Evo-1](https://hf.co/vrfai/evo1-libero-gguf) | Y | Y | Y | Y | Y |
332325
| [VLA-Adapter](https://hf.co/vrfai/vla-adapter-libero-gguf) | Y | Y | ~ | Y | Y |
333326
| [OpenVLA-OFT](https://hf.co/vrfai/openvla-oft-libero-gguf) | Y | Y | - | Y | ~ |
334-
| [VLA-JEPA](https://hf.co/vrfai/vla-jepa-libero) | Y | Y | - | Y | ~ |
327+
| [VLA-JEPA](https://hf.co/vrfai/vla-jepa-libero) | Y | Y | - | Y | Y |
335328

336329
---
337330

‎docs/backend/ov.md‎

Lines changed: 101 additions & 67 deletions
Original file line numberDiff line numberDiff line change
@@ -5,16 +5,15 @@ account of how far it currently runs. Like SYCL, OpenVINO is **not**
55
auto-detected: it needs an explicit `-DGGML_OPENVINO=ON` and the OpenVINO
66
runtime on the configure line.
77

8-
> **Status: five architectures translate faithfully, two drift, one is wrong.**
9-
> Evo-1 and VLA-Adapter agree with an F32 CPU reference to six decimal places;
10-
> SmolVLA and π0.5 to under 1e-3; GR00T N1.6 lands on the bar at 3.0e-3. VLA-JEPA (5.5e-3) and GR00T N1.5 (2.7e-2) run to
11-
> completion but sit outside the bar the SYCL backend is held to, for reasons only
12-
> partly established. GR00T N1.7 translates and returns plausible-looking actions
13-
> that are simply wrong. On the Arc B390 iGPU the speedup over the native CPU
14-
> backend runs from 3.1x to 8.2x. Eleven fixes were needed, nine of them inside
15-
> ggml's OpenVINO backend, which is written against llama.cpp's graphs and had
16-
> never seen a vision tower or an action expert - see
17-
> [What had to change](#what-had-to-change).
8+
> **Status: seven architectures translate faithfully, one is wrong.** SmolVLA,
9+
> π0.5, Evo-1, VLA-Adapter, GR00T N1.5, GR00T N1.6 and VLA-JEPA all agree with a
10+
> CPU-backend reference to 1.3e-3 or better on the OpenVINO CPU plugin - Evo-1 and
11+
> VLA-Adapter to about 3e-6. On the Arc B390 iGPU the speedup over the native CPU
12+
> backend runs from 3.0x to 9.6x. GR00T N1.7 translates and returns
13+
> plausible-looking actions that are simply wrong; it is the one open failure.
14+
> Twelve fixes were needed, ten of them inside ggml's OpenVINO backend, which is
15+
> written against llama.cpp's graphs and had never seen a vision tower or an
16+
> action expert - see [What had to change](#what-had-to-change).
1817
>
1918
> Note the baseline: OpenVINO executes the checkpoint's BF16 weights at F32, so
2019
> compare against `--weight-dtype f32` or you will charge the backend for a
@@ -175,17 +174,17 @@ reason in [Known issues](#known-issues).
175174

176175
| Model | input | CPU backend | OpenVINO CPU | OpenVINO GPU | OpenVINO NPU |
177176
|---|---|---:|---:|---:|---:|
178-
| SmolVLA | 512 | 1,364 ms | 1,340 ms | **446 ms** (3.1x) | 1,162 ms |
179-
| π0.5 | 224 | 2,802 ms | 4,285 ms | **641 ms** (4.4x) | 916 ms |
180-
| Evo-1 | 448 | 3,114 ms | 4,523 ms | **574 ms** (5.4x) | not supported |
181-
| VLA-Adapter | 224 | 1,228 ms | 1,603 ms | **162 ms** (7.6x) | not supported |
182-
| VLA-JEPA | 256 | 1,046 ms | 1,265 ms | **128 ms** (8.2x) | not supported |
183-
| GR00T N1.6 | 224 | 1,276 ms | not timed | **322 ms** (4.0x) | not supported |
184-
| GR00T N1.5 | 224 | not timed | not timed | not timed | not supported |
185-
186-
GR00T N1.5 is deliberately not timed: it drifts too far from the CPU backend to
187-
report a latency as if the two were doing the same work, and GR00T N1.7 is wrong
188-
outright.
177+
| VLA-JEPA | 256 | 1,046 ms | 1,265 ms | **127 ms** (8.2x) | returns NaN |
178+
| GR00T N1.5 | 224 | 1,420 ms | 2,199 ms | **148 ms** (9.6x) | plugin throws |
179+
| VLA-Adapter | 224 | 1,228 ms | 1,603 ms | **161 ms** (7.6x) | not supported |
180+
| GR00T N1.6 | 224 | 1,276 ms | 2,256 ms | **323 ms** (3.9x) | plugin throws |
181+
| SmolVLA | 512 | 1,364 ms | 1,340 ms | **451 ms** (3.0x) | 1,162 ms |
182+
| Evo-1 | 448 | 3,114 ms | 4,523 ms | **563 ms** (5.5x) | not supported |
183+
| π0.5 | 224 | 2,802 ms | 4,285 ms | **683 ms** (4.1x) | 916 ms |
184+
| GR00T N1.7 | 256 | not timed | not timed | not timed | not attempted |
185+
186+
GR00T N1.7 is not timed because it computes the wrong answer - a latency for work
187+
that is not the same work would be misleading.
189188

190189
The iGPU is the reason to use this backend, and it pays off most where the model
191190
is most vision-heavy. The OpenVINO CPU plugin is at best parity with ggml's own
@@ -209,22 +208,23 @@ The two references bracket the answer, and which one is tighter is arch-dependen
209208
- it turns on how much of a given checkpoint is BF16 in the first place. Report
210209
both and take the smaller as the fidelity figure:
211210

212-
| Model | vs BF16 reference (default) | vs F32 reference | what the gap was |
211+
| Model | vs BF16 reference | vs F32 reference | tighter reference |
213212
|---|---:|---:|---|
214-
| Evo-1 | 2.7e-3 | **3.6e-6** | almost entirely BF16 vs F32 |
215-
| VLA-Adapter | 2.1e-3 | **2.4e-6** | almost entirely BF16 vs F32 |
216-
| π0.5 | 8.9e-4 | **2.5e-4** | mostly |
217-
| SmolVLA | 1.2e-3 | **9.8e-4** | partly |
218-
| VLA-JEPA | 1.2e-2 | **5.5e-3** | about half |
219-
| GR00T N1.6 | **3.0e-3** | 4.7e-3 | none - BF16 is the tighter reference here |
220-
| GR00T N1.5 | 2.2e-2 | **2.7e-2** | none - not a precision effect |
213+
| VLA-Adapter | 2.1e-3 | **2.4e-6** | F32 |
214+
| Evo-1 | 2.7e-3 | **3.5e-6** | F32 |
215+
| π0.5 | 7.7e-4 | **3.5e-5** | F32 |
216+
| VLA-JEPA | 1.1e-2 | **1.1e-4** | F32 |
217+
| GR00T N1.5 | 5.5e-3 | **6.0e-4** | F32 |
218+
| GR00T N1.6 | 5.0e-3 | **1.0e-3** | F32 |
219+
| SmolVLA | **8.9e-4** | 1.3e-3 | BF16 |
220+
| GR00T N1.7 | 1.5e0 | 1.5e0 | neither - it is wrong |
221221

222222
Evo-1 and VLA-Adapter agree with an F32 reference to six decimal places, which is
223223
as close to "the translation is exact" as this harness can show - for those two,
224224
OpenVINO is doing F32 arithmetic and the BF16 comparison was measuring nothing but
225-
the dtype. GR00T N1.6 is the counterexample that stops this being a universal
226-
rule: it lands closer to the BF16 reference, so its checkpoint evidently is not
227-
uniformly BF16 where it matters. For context, the
225+
the dtype. SmolVLA is the counterexample that stops this being a universal rule: it lands
226+
closer to the BF16 reference, so its checkpoint evidently is not uniformly BF16
227+
where it matters. For context, the
228228
CPU backend's own output moves by 2.0e-3 (SmolVLA), 2.7e-3 (Evo-1) or 1.1e-2
229229
(VLA-JEPA) when you flip that one flag, so the model's intrinsic sensitivity to
230230
precision is the same size as the numbers being reported.
@@ -235,29 +235,27 @@ Against the F32 reference, which is the fidelity number:
235235

236236
| Model | OpenVINO CPU | OpenVINO GPU | OpenVINO NPU |
237237
|---|---:|---:|---:|
238-
| SmolVLA | 9.8e-4 | 1.1e-3 | 1.6e-2 |
239-
| π0.5 | 2.5e-4 | 5.0e-4 | 1.6e-3 |
240-
| Evo-1 | 3.6e-6 | 1.2e-3 | not supported |
241-
| VLA-Adapter | 2.4e-6 | 2.7e-3 | not supported |
242-
| GR00T N1.6 | 3.0e-3 | 3.3e-3 | plugin throws |
243-
| VLA-JEPA | 5.5e-3 | 5.4e-3 | returns NaN |
244-
| GR00T N1.5 | 2.7e-2 | 6.1e-2 | NPUW throws |
238+
| VLA-Adapter | 2.4e-6 | 3.2e-3 | not supported |
239+
| Evo-1 | 3.5e-6 | 1.2e-3 | not supported |
240+
| π0.5 | 3.5e-5 | 4.7e-4 | 9.9e-4 |
241+
| VLA-JEPA | 1.1e-4 | 8.7e-3 | returns NaN |
242+
| GR00T N1.5 | 6.0e-4 | 4.7e-3 | plugin throws |
243+
| GR00T N1.6 | 1.0e-3 | 2.0e-3 | plugin throws |
244+
| SmolVLA | 8.9e-4 | 1.3e-3 | 1.7e-2 |
245245
| GR00T N1.7 | 1.5e0 | 1.5e0 | not attempted |
246246

247-
GR00T N1.6 sits right on the bar (3.0e-3 against 2.9e-3) rather than comfortably
248-
inside it.
247+
Every arch except GR00T N1.7 is inside the 2.9e-3 bar on the CPU plugin, most by
248+
one to three orders of magnitude. On the GPU the picture is looser because that
249+
plugin computes in F16: VLA-JEPA (8.7e-3), GR00T N1.5 (4.7e-3) and VLA-Adapter (3.2e-3) sit outside the bar there
250+
even though all three are far inside it on the CPU plugin. Judge translation
251+
fidelity on the CPU plugin; treat the GPU as a separate precision target.
249252

250-
The GPU is consistently looser than the CPU plugin because that plugin runs F16
251-
internally; it stays within the band the SYCL backend is held to (2.9e-3) for
252-
every arch except the two below.
253+
Two effects explain the residuals that remain. The GPU plugin's F16 arithmetic is
254+
one. The other is SmolVLA on the NPU (1.7e-2), whose compile config turns on
255+
dynamic quantization - π0.5 on the same device stays at 9.9e-4, so that is a
256+
property of the model on that device rather than of the backend.
253257

254-
**Two archs remain unexplained.** VLA-JEPA sits at 5.5e-3, roughly 2x the bar,
255-
with half its original gap accounted for by precision and half not. GR00T N1.5 is
256-
the real outlier: 2.7e-2 that does **not** move with the weight dtype, does not
257-
improve with `--mm-prec f32`, and gets worse on the GPU (6.1e-2). It does not use
258-
mrope or flash attention, so those paths are ruled out. Both stay `~` in the
259-
README matrix. GR00T N1.7 is wrong outright - see
260-
[What is left](#what-is-left).
258+
GR00T N1.7 is wrong outright - see [What is left](#what-is-left).
261259

262260
## What had to change
263261

@@ -287,13 +285,14 @@ larger through a model builder that assumes a decoder-only LLM. The literal path
287285
is the one that fits a vision tower and an action expert. An explicit setting
288286
still wins.
289287

290-
The other seven are in ggml's OpenVINO backend itself, applied by
288+
The other ten are in ggml's OpenVINO backend itself, applied by
291289
`scripts/patch_ggml_openvino.py` at configure time. Its docstring carries the
292290
detail; in short each narrows an llama.cpp-shaped assumption that is stricter
293291
than the ggml contract, or fills a gap:
294292

295293
| Fix | What it addresses |
296294
|---|---|
295+
| **GELU translated as tanh, not erf** | **assumes ggml's GELU is the exact erf form** |
297296
| Intel OpenCL platform selection | assumes the first OpenCL platform is Intel's |
298297
| RESHAPE `op_case` guard | assumes a reshape flattening dims 0-2 is the KV-cache flatten |
299298
| SDPA K/V converted with Q | assumes K/V arrive as F16 because the KV cache is |
@@ -303,7 +302,17 @@ than the ggml contract, or fills a gap:
303302
| Missing op translators | RELU, GELU_ERF, NEG, SQR had no table entry |
304303
| Naive-path graph cache | that path re-compiled the whole model on every graph_compute |
305304

306-
Three are worth expanding.
305+
Four are worth expanding.
306+
307+
**The GELU mode** is the highest-yield single fix in the list. ggml's
308+
`GGML_UNARY_OP_GELU` is the *tanh* approximation - its CPU kernel additionally
309+
reads an fp16 lookup table - while `ov::op::v7::Gelu` defaults to the exact erf
310+
formulation, and both ggml GELU ops were mapped onto that default. The error per
311+
node is small, but a Qwen3-VL vision tower contains dozens of them and it
312+
compounds through the encoder. Setting the mode explicitly moved VLA-JEPA from
313+
5.5e-3 to 1.1e-4 (48x) and GR00T N1.5 from 2.7e-2 to 6.0e-4 (45x), turning both
314+
from "runs but drifts" into supported, and improved GR00T N1.6 and π0.5 too. It
315+
is worth upstreaming alongside the position-input fix.
307316

308317
**Position inputs** is what carries an arch through to a full prediction. Every
309318
tensor feeding a `GGML_OP_ROPE`'s second input was renamed to a single parameter
@@ -357,7 +366,7 @@ variable is set. Unverified guess at the cause: the blob key does not capture
357366
something that differs between vla.cpp's several graphs, so one graph gets
358367
another's blob. In practice, pay the compile once per process and leave it unset.
359368

360-
**The NPU takes two of the five archs, and fails three different ways.** SmolVLA
369+
**The NPU takes two of the eight archs, and fails four different ways.** SmolVLA
361370
and π0.5 run. The others do not:
362371

363372
| Model | NPU outcome |
@@ -410,22 +419,47 @@ vision tower and the same interleaved mrope and is only ~1% off, which points at
410419
GR00T N1.7's own action expert rather than anything shared. Not yet diagnosed;
411420
the arch stays `-` in the README matrix.
412421

413-
**VLA-JEPA and GR00T N1.5 drift, and only half of it is explained.** Measured
414-
against an F32 reference, VLA-JEPA is 5.5e-3 and GR00T N1.5 is 2.7e-2, against a
415-
bar of 2.9e-3. Half of VLA-JEPA's original gap turned out to be the BF16/F32
416-
baseline; the rest did not. GR00T N1.5's did not move at all with the weight
417-
dtype, nor with `--mm-prec f32`, and it is worse on the GPU (6.1e-2) than the CPU
418-
plugin. Neither arch uses flash attention, and GR00T N1.5 does not use mrope, so
419-
those paths are ruled out. Both stay `~` rather than `Y` until the residual is
420-
understood - the next thing to try is bisecting the graph, since vla.cpp runs the
421-
vision tower and the action head as separate `ggml_backend_graph_compute` calls
422-
and each could be compared in isolation.
422+
**GR00T N1.7 is the one open failure.** It translates, returns a full action
423+
chunk, and the values are wrong: max|delta| 1.477, rms 8.3e-2 against a peak of
424+
0.948, with 86% of 5280 values off by more than 1e-2. Deterministic, and
425+
identical on the CPU and GPU plugins. Ruled out so far, each with evidence:
426+
427+
- *Precision.* Identical against BF16 and F32 references. OpenVINO is 1530x less
428+
weight-dtype-sensitive than ggml CPU here, and the model's own bf16/f32
429+
sensitivity is rms 7.2e-4 against the error's 8.3e-2.
430+
- *Translation-path selection.* The naive path is taken by default and
431+
`is_model_splitted` returns false for every graph of this arch; forcing all
432+
graphs down the decoder-only-LLM path instead changes the answer by 1e-5 while
433+
both remain 1.477 from the reference.
434+
- *Sequence length and token composition.* Swept 3.7x (SEQ 70 to 262) by two
435+
independent routes; relative error stayed within 0.2545-0.2841 and the fraction
436+
off by >1e-2 within 84.7-86.6%. Nothing accumulates.
437+
- *Matmul precision and the compiled-model cache.* Both bitwise no-ops.
438+
- *The GELU mode*, which fixed VLA-JEPA and GR00T N1.5, moves N1.7 by nothing.
439+
- *A missing or unsupported op.* GR00T N1.6 (works) and N1.7 (broken) have the
440+
same 17-op vocabulary, and LM layer 0 is node-for-node identical except that
441+
N1.7's position input is `[4*SEQ]` for IMROPE where N1.6's is `[SEQ]` for NEOX.
442+
VLA-JEPA also uses IMROPE and is now clean, which weakens that lead.
443+
444+
Three earlier claims about N1.7 turned out to be **measurement artifacts** and
445+
should not be reused. That its first LM block is 73-95% wrong: OpenVINO's
446+
`lm_h_00..03` dumps are bit-exact copies of the *input* arrays and `lm_h_04..15`
447+
are zeros, because the main graph has exactly one real output, `action_pred`.
448+
That `--flash-attn 1` yields 0.948: it actually fails with "Got less inputs than
449+
expected" and returns no actions, and 0.948 is `max|reference|` - the number you
450+
get comparing against nothing. And that honouring `GGML_TENSOR_FLAG_OUTPUT`
451+
changes the answer: same artifact.
452+
453+
The next step follows from the first artifact. The stage dump cannot see inside
454+
the graph because ggml-openvino writes back only true graph outputs, so either
455+
add a debug mode that materialises selected intermediates as `ov::Result`s, or
456+
bisect with cut-down graphs.
423457

424458
**Untested archs.** π0 and OpenVLA-OFT are untested here for want of a local
425459
checkpoint, not because anything is known to block them; both use op sets already
426-
covered by tested archs (π0 matches π0.5, OpenVLA-OFT matches VLA-Adapter). Treat
427-
any untested arch as `-` in the README matrix until it has actually produced
428-
actions.
460+
covered by tested archs (π0 matches π0.5, OpenVLA-OFT matches VLA-Adapter). Both
461+
are `~` in the README matrix; treat any untested arch that way until it has
462+
actually produced actions.
429463

430464
**Splitting across devices.** Intel's own
431465
[π0.5 write-up](https://docs.openedgeplatform.intel.com/2026.1/OEP-articles/publications/optimizing-pi0.5-lva-model.html)

0 commit comments

Comments
 (0)