Repository navigation
0.3.53: ship one GPU implementation — publish backend-cuda, remove backend-tornado - #239
Merged
Merged
Conversation
Both device gates pass on an RTX 4090 at compute capability 8.9, Granite 4.1 3B Q4_K_M, 20 prompts, models revision 15661fb: G1 token parity PASSED 1280 token ids identical, firstDivergence null G4 decode speed PASSED 6.184x against a 3.00x gate, 31.26 vs 5.06 tok/s G1's parity.selfTest is false, which is the field that makes it G1 evidence at all rather than a CPU-vs-CPU self-test, and routing.totalDeclinedProjections is 0, so those 1280 tokens were produced with the projections actually on the device and not by a quiet fallback that would have made parity trivial. backend-cuda joins the publication allowlist. The artifact is opt-in and nothing activates it by accident: no META-INF/services entry so no ServiceLoader discovers it, a consumer must call CudaGgufBatchedMatrixKernel.open() and inject the kernel, and -Dmodels.cuda.disabled=true is a kill switch on top. One sm_80 PTX module serves every device of capability 8.0 or above, carried under META-INF/models/cuda/ with a SHA-256 the loader recomputes. Adding it to the allowlist turned on the published-module coverage gate, which failed at 0.28 against 0.80. That is not a formality. CudaDriver is the FFM binding to libcuda.so.1 -- every method is a downcall -- and CudaGgufBatchedMatrixKernel is the dispatch path behind it; together they are 2,296 of the module's 2,491 missed instructions and cannot execute on a host with no driver. Those two classes are exempted, with the reasoning in backend-cuda/build.gradle.kts, and the 0.80 bar still applies to everything a host can reach -- CudaRoutingCounters already measures 391 of 403 instructions and Q8KActivations 177 of 193. Lowering the global minimum to admit this module would have hidden a real gap in every other published module. No other published module changed: v0.3.52..HEAD touches none of them, so the rest of the release is byte-identical and the changelog says so. What this does not establish, and the records say so rather than implying otherwise: one host and one model. G1 carries no tolerance, so each qualifying hardware profile and each architecture family earns its own run -- this was cc 8.9 where the September failure was an A40 at 8.6, and Granite 4.1 3B exercises neither a routed mixture-of-experts FFN nor a per-layer feed-forward width. Still outstanding before Actions -> Release: RELEASING.md requires the public loader and exact CPU/GPU parity gate on each qualified NVIDIA profile, retained under models-accelerator-bench/results/. That is backend-tornado, which is published and needs TornadoVM installed to exercise, and it has not been run. spotlessCheck build green: 470 classes, 2596 tests, zero failures.
The agent worktrees under .claude/worktrees are embedded git repositories, so adding them produces a commit that clones cannot resolve. Caught during 0.3.53 preparation and removed from that commit; ignored here so it cannot recur.
… launches
RELEASING.md requires the public loader and exact CPU/GPU parity gate on each
qualified NVIDIA profile, retained under models-accelerator-bench/results/.
This is that run on an RTX 4090 with TornadoVM v5.2.0-jdk25 built from source
for the PTX backend, and it is retained as a failed precondition rather than a
pass.
The gate exited 0 and reported accelerated=true device=cuda-0 reason=eligible,
readiness 32,171 ms, plans=321, prefill 69.34 tok/s, decode 6.49 tok/s, and
routed projections {Q4_K=16800, Q6_K=2870}. Both format counts being non-zero
is the check the models-bench README asks for, so the mixed-format partial
routing trap is not what happened.
What the report does not contain is 36 occurrences of
[TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701
CUDA 701 is CUDA_ERROR_LAUNCH_OUT_OF_RESOURCES. Thirty-six launches failed on
the device, and the JSON has no field for a failed launch, so a reader of the
artifact alone sees accelerated=true and 6.49 tok/s. Exit status is therefore
not evidence the device path worked, and the throughput is not filed as a
measurement of accelerated decode.
The number supports that reading. 6.49 tok/s is barely above the 5.06 tok/s
Vector API control measured on the same model and GPU class in
benchmark-results/2026-10-07-g4-dualpath, where backend-cuda reached 31.26.
Readiness of 32,171 ms is inside G5's 120,000 ms ceiling and two orders of
magnitude above backend-cuda's 344-353 ms on the same hardware.
This is the third instance in one day of evidence reporting success while
something silently did not run: backend-cuda left 53.3% of a layer's projection
arithmetic on the CPU behind an empty refusals map; the counter added to catch
that printed None because it was never serialised; and here a gate returns 0
over 36 failed launches. The fix is the same shape each time -- the artifact has
to carry the negative. accelerator-profile needs a launch-failure count, and
--require true should fail on a nonzero one.
Also recorded for whoever runs this next, because each cost a pod: the pinned
tag builds with --jdk jdk25, which yields -Pjdk25,ptx-backend, and the wrong
value compiles at -source 8 and dies on sealed classes; the assembled SDK lands
under dist/tornadovm-<ver>-<backend>-<platform>/tornadovm-<ver>-<backend>/ and
TORNADOVM_HOME must point there because bin/tornado reads
$TORNADOVM_HOME/etc/tornado.backend; and :models-bench:installDist pulls in
:backend-cuda:compilePtx, so cargo and the pinned nightly's rust-src are
required.
…er one
Measured on the same RTX 4090 and the same model on the same day:
backend-cuda backend-tornado
G1 exact parity 1280/1280 not reached
decode 31.26 tok/s 6.49 tok/s
CPU control 5.06 tok/s 5.06 tok/s
readiness 344 ms 32,171 ms
device launch errors 0 36 x CUDA 701
self-contained yes no
A second implementation of the same capability, five times slower than the
first and barely ahead of the CPU path it exists to accelerate, is not worth
the surface it costs. The 36 cuLaunchKernel failures are
CUDA_ERROR_LAUNCH_OUT_OF_RESOURCES and its report had no field to record them,
so the gate returned 0 with accelerated=true over a device that was erroring.
It could not be self-contained either: a Maven artifact cannot carry TornadoVM's
device runtime, so an application had to install a matching distribution and
launch through its own launcher, which defeats the point of adding a dependency.
Removed with the module, because each existed only to serve it: the
accelerator-profile command and its test, the TornadoVM benchmark arm in
models-accelerator-bench with its orphaned tests, the tornado-api and
tornado-runtime dependencies, and the RELEASING.md precondition requiring a
loader and parity gate on each NVIDIA profile under
models-accelerator-bench/results/. That precondition existed for
backend-tornado; backend-cuda's gates are cuda-kernel-gate
--mode capability|parity|decode and need no external runtime.
Docs are rewritten rather than left stale: README, modules.adoc, index.adoc,
architecture.adoc, models-accelerator-bench/README and gpu-acceleration.adoc,
which was 188 lines of TornadoVM and is now 107 documenting backend-cuda with
its measured gates.
Deliberately kept: the KV-ridge experiment that happened to live in
models-accelerator-bench, which is unrelated work, and the August measurement
records that name backend-tornado, which are what was measured then and are not
rewritten to match a later decision.
The 0.3.53 notes said no published module's behaviour changed. That was true
when written and is not now, so it is corrected: the published surface changes
in exactly two ways, backend-cuda added and backend-tornado removed.
clean spotlessCheck build: 466 classes, 2587 tests, zero failures.
…ties
CI failed on :verifyReleaseMetadata with "Antora display_version must match
0.3.53". The version bump had changed gradle.properties alone, and that task
exists because a release where the published coordinate and the user-facing
version disagree is a release that documents the wrong thing.
Running it locally then surfaced two more classes it checks, in order:
Antora display_version docs/content/antora.yml
models-version attribute docs/content/antora.yml
docs package version docs/package.json, docs/package-lock.json
landing page pill and snippet docs/landing/index.html
native crate version model-kernels Cargo.toml and Cargo.lock
published coordinates in prose models-backend-apple/README.md,
models-rag/README.md
notebook release defaults notebooks/.env.example,
notebooks/docker-compose.yml,
notebooks/jupyter/prepare-classpath.sh,
notebooks/README.md
The Cargo.lock edit is anchored on the jmodels-kernels package entry rather
than applied globally, so a dependency that happens to carry the same version
string is not rewritten.
The cause of the miss is worth recording: the local gate was spotlessCheck
build -x :docs:build, and verifyReleaseMetadata hangs off complianceCheck,
which the release workflow runs and that command does not. complianceCheck now
passes locally, which is the check that should have been run before pushing a
version bump.
The merge-and-release chain stopped on the failing check instead of merging,
which is what it was gated to do.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Models ships one implementation of a given thing, and for GPU acceleration that is
backend-cuda. Measured on the same RTX 4090 and the same model on the same day:backend-cudabackend-tornadoPublished: backend-cuda
Both device gates pass. G1 — 1280 of 1280 token ids identical,
firstDivergence: null,selfTest: false, andtotalDeclinedProjections: 0so the parity was earned with projections genuinely on the device. G4 — 6.184x against a 3.00x gate.Opt-in by construction: no
META-INF/services, so noServiceLoaderfinds it; a consumer callsCudaGgufBatchedMatrixKernel.open(), with-Dmodels.cuda.disabled=trueas a kill switch. Onesm_80PTX module inside the jar serves every device of capability 8.0+.Adding it activated the published-module coverage gate, which failed at 0.28 vs 0.80.
CudaDriveris the FFM binding tolibcuda— every method a downcall — and with the dispatch path it is 2,296 of 2,491 missed instructions, uncoverable without a driver. Those two classes are exempted with the reasoning in the build file; the 0.80 bar still applies to everything host-reachable (CudaRoutingCounters391/403,Q8KActivations177/193). Lowering the global minimum would have hidden real gaps in every other published module.Removed: backend-tornado
The 36 failures are
CUDA_ERROR_LAUNCH_OUT_OF_RESOURCES, and its report had no field to record them — the gate returned 0 withaccelerated=trueover a device that was erroring. It also cannot be self-contained: a Maven artifact cannot carry TornadoVM's device runtime, so an application had to install a matching distribution and launch through its own launcher.Removed with it, each existing only to serve it: the
accelerator-profilecommand and test, the TornadoVM benchmark arm and its orphaned tests, thetornado-api/tornado-runtimedependencies, and theRELEASING.mdprecondition requiring a loader and parity gate on each NVIDIA profile — that requirement existed forbackend-tornado, so the release blocker goes with the module.Docs rewritten rather than left stale, including
gpu-acceleration.adoc: 188 lines of TornadoVM, now 107 documentingbackend-cudawith its gates.Deliberately kept: the unrelated KV-ridge experiment that shared the module, and the August measurement records naming
backend-tornado— those are what was measured then and are not rewritten to match a later decision.Also in here
The FFN gate/up projections never reached the device (
multiplyDualunimplemented,isDualEligibleinheriting SPIfalse) — 53.3% of a layer's projection arithmetic on the CPU, which is what took G4 from 1.788x to 6.184x. PlusCudaRoutingCounters.declined(...)with the shape in the key, now serialised into the report, because an unimplemented path is not a refusal and an empty refusals map read as "everything ran on the device".One host and one model. G1 has no tolerance, so each qualifying hardware profile and architecture family earns its own run. The published surface changes in exactly two ways:
backend-cudaadded,backend-tornadoremoved.clean
spotlessCheck build: 466 classes, 2587 tests, zero failures.