Skip to content

0.3.53: ship one GPU implementation — publish backend-cuda, remove backend-tornado - #239

Merged
bsbodden merged 5 commits into
mainfrom
feat/publish-backend-cuda
Oct 7, 2026
Merged

bsbodden merged 5 commits into
mainfrom
feat/publish-backend-cuda

Conversation

@bsbodden

@bsbodden bsbodden commented Oct 7, 2026 •

Copy link
Copy Markdown
Member

Models ships one implementation of a given thing, and for GPU acceleration that is backend-cuda. Measured on the same RTX 4090 and the same model on the same day:

backend-cuda backend-tornado
G1 exact token parity 1280 / 1280 identical not reached
decode 31.26 tok/s 6.49 tok/s
CPU control, same host 5.06 tok/s 5.06 tok/s
readiness 344 ms 32,171 ms
device launch errors 0 36 × CUDA 701
self-contained artifact yes no

Published: backend-cuda

Both device gates pass. G1 — 1280 of 1280 token ids identical, firstDivergence: null, selfTest: false, and totalDeclinedProjections: 0 so the parity was earned with projections genuinely on the device. G4 — 6.184x against a 3.00x gate.

Opt-in by construction: no META-INF/services, so no ServiceLoader finds it; a consumer calls CudaGgufBatchedMatrixKernel.open(), with -Dmodels.cuda.disabled=true as a kill switch. One sm_80 PTX module inside the jar serves every device of capability 8.0+.

Adding it activated the published-module coverage gate, which failed at 0.28 vs 0.80. CudaDriver is the FFM binding to libcuda — every method a downcall — and with the dispatch path it is 2,296 of 2,491 missed instructions, uncoverable without a driver. Those two classes are exempted with the reasoning in the build file; the 0.80 bar still applies to everything host-reachable (CudaRoutingCounters 391/403, Q8KActivations 177/193). Lowering the global minimum would have hidden real gaps in every other published module.

Removed: backend-tornado

The 36 failures are CUDA_ERROR_LAUNCH_OUT_OF_RESOURCES, and its report had no field to record them — the gate returned 0 with accelerated=true over a device that was erroring. It also cannot be self-contained: a Maven artifact cannot carry TornadoVM's device runtime, so an application had to install a matching distribution and launch through its own launcher.

Removed with it, each existing only to serve it: the accelerator-profile command and test, the TornadoVM benchmark arm and its orphaned tests, the tornado-api/tornado-runtime dependencies, and the RELEASING.md precondition requiring a loader and parity gate on each NVIDIA profile — that requirement existed for backend-tornado, so the release blocker goes with the module.

Docs rewritten rather than left stale, including gpu-acceleration.adoc: 188 lines of TornadoVM, now 107 documenting backend-cuda with its gates.

Deliberately kept: the unrelated KV-ridge experiment that shared the module, and the August measurement records naming backend-tornado — those are what was measured then and are not rewritten to match a later decision.

Also in here

The FFN gate/up projections never reached the device (multiplyDual unimplemented, isDualEligible inheriting SPI false) — 53.3% of a layer's projection arithmetic on the CPU, which is what took G4 from 1.788x to 6.184x. Plus CudaRoutingCounters.declined(...) with the shape in the key, now serialised into the report, because an unimplemented path is not a refusal and an empty refusals map read as "everything ran on the device".

One host and one model. G1 has no tolerance, so each qualifying hardware profile and architecture family earns its own run. The published surface changes in exactly two ways: backend-cuda added, backend-tornado removed.

clean spotlessCheck build: 466 classes, 2587 tests, zero failures.

Both device gates pass on an RTX 4090 at compute capability 8.9, Granite 4.1 3B
Q4_K_M, 20 prompts, models revision 15661fb:

  G1 token parity   PASSED   1280 token ids identical, firstDivergence null
  G4 decode speed   PASSED   6.184x against a 3.00x gate, 31.26 vs 5.06 tok/s

G1's parity.selfTest is false, which is the field that makes it G1 evidence at
all rather than a CPU-vs-CPU self-test, and routing.totalDeclinedProjections is
0, so those 1280 tokens were produced with the projections actually on the
device and not by a quiet fallback that would have made parity trivial.

backend-cuda joins the publication allowlist. The artifact is opt-in and
nothing activates it by accident: no META-INF/services entry so no ServiceLoader
discovers it, a consumer must call CudaGgufBatchedMatrixKernel.open() and inject
the kernel, and -Dmodels.cuda.disabled=true is a kill switch on top. One sm_80
PTX module serves every device of capability 8.0 or above, carried under
META-INF/models/cuda/ with a SHA-256 the loader recomputes.

Adding it to the allowlist turned on the published-module coverage gate, which
failed at 0.28 against 0.80. That is not a formality. CudaDriver is the FFM
binding to libcuda.so.1 -- every method is a downcall -- and
CudaGgufBatchedMatrixKernel is the dispatch path behind it; together they are
2,296 of the module's 2,491 missed instructions and cannot execute on a host
with no driver. Those two classes are exempted, with the reasoning in
backend-cuda/build.gradle.kts, and the 0.80 bar still applies to everything a
host can reach -- CudaRoutingCounters already measures 391 of 403 instructions
and Q8KActivations 177 of 193. Lowering the global minimum to admit this module
would have hidden a real gap in every other published module.

No other published module changed: v0.3.52..HEAD touches none of them, so the
rest of the release is byte-identical and the changelog says so.

What this does not establish, and the records say so rather than implying
otherwise: one host and one model. G1 carries no tolerance, so each qualifying
hardware profile and each architecture family earns its own run -- this was cc
8.9 where the September failure was an A40 at 8.6, and Granite 4.1 3B exercises
neither a routed mixture-of-experts FFN nor a per-layer feed-forward width.

Still outstanding before Actions -> Release: RELEASING.md requires the public
loader and exact CPU/GPU parity gate on each qualified NVIDIA profile, retained
under models-accelerator-bench/results/. That is backend-tornado, which is
published and needs TornadoVM installed to exercise, and it has not been run.

spotlessCheck build green: 470 classes, 2596 tests, zero failures.
The agent worktrees under .claude/worktrees are embedded git repositories, so
adding them produces a commit that clones cannot resolve. Caught during 0.3.53
preparation and removed from that commit; ignored here so it cannot recur.
… launches

RELEASING.md requires the public loader and exact CPU/GPU parity gate on each
qualified NVIDIA profile, retained under models-accelerator-bench/results/.
This is that run on an RTX 4090 with TornadoVM v5.2.0-jdk25 built from source
for the PTX backend, and it is retained as a failed precondition rather than a
pass.

The gate exited 0 and reported accelerated=true device=cuda-0 reason=eligible,
readiness 32,171 ms, plans=321, prefill 69.34 tok/s, decode 6.49 tok/s, and
routed projections {Q4_K=16800, Q6_K=2870}. Both format counts being non-zero
is the check the models-bench README asks for, so the mixed-format partial
routing trap is not what happened.

What the report does not contain is 36 occurrences of

  [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701

CUDA 701 is CUDA_ERROR_LAUNCH_OUT_OF_RESOURCES. Thirty-six launches failed on
the device, and the JSON has no field for a failed launch, so a reader of the
artifact alone sees accelerated=true and 6.49 tok/s. Exit status is therefore
not evidence the device path worked, and the throughput is not filed as a
measurement of accelerated decode.

The number supports that reading. 6.49 tok/s is barely above the 5.06 tok/s
Vector API control measured on the same model and GPU class in
benchmark-results/2026-10-07-g4-dualpath, where backend-cuda reached 31.26.
Readiness of 32,171 ms is inside G5's 120,000 ms ceiling and two orders of
magnitude above backend-cuda's 344-353 ms on the same hardware.

This is the third instance in one day of evidence reporting success while
something silently did not run: backend-cuda left 53.3% of a layer's projection
arithmetic on the CPU behind an empty refusals map; the counter added to catch
that printed None because it was never serialised; and here a gate returns 0
over 36 failed launches. The fix is the same shape each time -- the artifact has
to carry the negative. accelerator-profile needs a launch-failure count, and
--require true should fail on a nonzero one.

Also recorded for whoever runs this next, because each cost a pod: the pinned
tag builds with --jdk jdk25, which yields -Pjdk25,ptx-backend, and the wrong
value compiles at -source 8 and dies on sealed classes; the assembled SDK lands
under dist/tornadovm-<ver>-<backend>-<platform>/tornadovm-<ver>-<backend>/ and
TORNADOVM_HOME must point there because bin/tornado reads
$TORNADOVM_HOME/etc/tornado.backend; and :models-bench:installDist pulls in
:backend-cuda:compilePtx, so cargo and the pinned nightly's rust-src are
required.
…er one

Measured on the same RTX 4090 and the same model on the same day:

                        backend-cuda      backend-tornado
  G1 exact parity       1280/1280         not reached
  decode                31.26 tok/s       6.49 tok/s
  CPU control           5.06 tok/s        5.06 tok/s
  readiness             344 ms            32,171 ms
  device launch errors  0                 36 x CUDA 701
  self-contained        yes               no

A second implementation of the same capability, five times slower than the
first and barely ahead of the CPU path it exists to accelerate, is not worth
the surface it costs. The 36 cuLaunchKernel failures are
CUDA_ERROR_LAUNCH_OUT_OF_RESOURCES and its report had no field to record them,
so the gate returned 0 with accelerated=true over a device that was erroring.
It could not be self-contained either: a Maven artifact cannot carry TornadoVM's
device runtime, so an application had to install a matching distribution and
launch through its own launcher, which defeats the point of adding a dependency.

Removed with the module, because each existed only to serve it: the
accelerator-profile command and its test, the TornadoVM benchmark arm in
models-accelerator-bench with its orphaned tests, the tornado-api and
tornado-runtime dependencies, and the RELEASING.md precondition requiring a
loader and parity gate on each NVIDIA profile under
models-accelerator-bench/results/. That precondition existed for
backend-tornado; backend-cuda's gates are cuda-kernel-gate
--mode capability|parity|decode and need no external runtime.

Docs are rewritten rather than left stale: README, modules.adoc, index.adoc,
architecture.adoc, models-accelerator-bench/README and gpu-acceleration.adoc,
which was 188 lines of TornadoVM and is now 107 documenting backend-cuda with
its measured gates.

Deliberately kept: the KV-ridge experiment that happened to live in
models-accelerator-bench, which is unrelated work, and the August measurement
records that name backend-tornado, which are what was measured then and are not
rewritten to match a later decision.

The 0.3.53 notes said no published module's behaviour changed. That was true
when written and is not now, so it is corrected: the published surface changes
in exactly two ways, backend-cuda added and backend-tornado removed.

clean spotlessCheck build: 466 classes, 2587 tests, zero failures.
@bsbodden bsbodden changed the title Release 0.3.53: publish backend-cuda, with G1 and G4 both passing 0.3.53: ship one GPU implementation — publish backend-cuda, remove backend-tornado Oct 7, 2026
…ties

CI failed on :verifyReleaseMetadata with "Antora display_version must match
0.3.53". The version bump had changed gradle.properties alone, and that task
exists because a release where the published coordinate and the user-facing
version disagree is a release that documents the wrong thing.

Running it locally then surfaced two more classes it checks, in order:

  Antora display_version          docs/content/antora.yml
  models-version attribute        docs/content/antora.yml
  docs package version            docs/package.json, docs/package-lock.json
  landing page pill and snippet   docs/landing/index.html
  native crate version            model-kernels Cargo.toml and Cargo.lock
  published coordinates in prose  models-backend-apple/README.md,
                                  models-rag/README.md
  notebook release defaults       notebooks/.env.example,
                                  notebooks/docker-compose.yml,
                                  notebooks/jupyter/prepare-classpath.sh,
                                  notebooks/README.md

The Cargo.lock edit is anchored on the jmodels-kernels package entry rather
than applied globally, so a dependency that happens to carry the same version
string is not rewritten.

The cause of the miss is worth recording: the local gate was spotlessCheck
build -x :docs:build, and verifyReleaseMetadata hangs off complianceCheck,
which the release workflow runs and that command does not. complianceCheck now
passes locally, which is the check that should have been run before pushing a
version bump.

The merge-and-release chain stopped on the failing check instead of merging,
which is what it was gated to do.
@bsbodden
bsbodden merged commit 70999a3 into main Oct 7, 2026
31 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant