From c2ed77c7a6fb40bf202523e9683670ca3e127fba Mon Sep 17 00:00:00 2001 From: Brian Sam-Bodden Date: Wed, 7 Oct 2026 12:38:02 -0700 Subject: [PATCH 1/5] Release 0.3.53: publish backend-cuda, with G1 and G4 both passing Both device gates pass on an RTX 4090 at compute capability 8.9, Granite 4.1 3B Q4_K_M, 20 prompts, models revision 15661fb2: G1 token parity PASSED 1280 token ids identical, firstDivergence null G4 decode speed PASSED 6.184x against a 3.00x gate, 31.26 vs 5.06 tok/s G1's parity.selfTest is false, which is the field that makes it G1 evidence at all rather than a CPU-vs-CPU self-test, and routing.totalDeclinedProjections is 0, so those 1280 tokens were produced with the projections actually on the device and not by a quiet fallback that would have made parity trivial. backend-cuda joins the publication allowlist. The artifact is opt-in and nothing activates it by accident: no META-INF/services entry so no ServiceLoader discovers it, a consumer must call CudaGgufBatchedMatrixKernel.open() and inject the kernel, and -Dmodels.cuda.disabled=true is a kill switch on top. One sm_80 PTX module serves every device of capability 8.0 or above, carried under META-INF/models/cuda/ with a SHA-256 the loader recomputes. Adding it to the allowlist turned on the published-module coverage gate, which failed at 0.28 against 0.80. That is not a formality. CudaDriver is the FFM binding to libcuda.so.1 -- every method is a downcall -- and CudaGgufBatchedMatrixKernel is the dispatch path behind it; together they are 2,296 of the module's 2,491 missed instructions and cannot execute on a host with no driver. Those two classes are exempted, with the reasoning in backend-cuda/build.gradle.kts, and the 0.80 bar still applies to everything a host can reach -- CudaRoutingCounters already measures 391 of 403 instructions and Q8KActivations 177 of 193. Lowering the global minimum to admit this module would have hidden a real gap in every other published module. No other published module changed: v0.3.52..HEAD touches none of them, so the rest of the release is byte-identical and the changelog says so. What this does not establish, and the records say so rather than implying otherwise: one host and one model. G1 carries no tolerance, so each qualifying hardware profile and each architecture family earns its own run -- this was cc 8.9 where the September failure was an A40 at 8.6, and Granite 4.1 3B exercises neither a routed mixture-of-experts FFN nor a per-layer feed-forward width. Still outstanding before Actions -> Release: RELEASING.md requires the public loader and exact CPU/GPU parity gate on each qualified NVIDIA profile, retained under models-accelerator-bench/results/. That is backend-tornado, which is published and needs TornadoVM installed to exercise, and it has not been run. spotlessCheck build green: 470 classes, 2596 tests, zero failures. --- CHANGELOG.md | 45 +++++++ RELEASING.md | 15 ++- backend-cuda/build.gradle.kts | 29 +++++ .../2026-10-07-g1-parity/README.md | 60 +++++++++ .../2026-10-07-g1-parity/rust-ptx-parity.json | 119 ++++++++++++++++++ .../2026-10-07-g1-parity/worker.log | 71 +++++++++++ build.gradle.kts | 1 + gradle.properties | 2 +- 8 files changed, 340 insertions(+), 2 deletions(-) create mode 100644 benchmark-results/2026-10-07-g1-parity/README.md create mode 100644 benchmark-results/2026-10-07-g1-parity/rust-ptx-parity.json create mode 100644 benchmark-results/2026-10-07-g1-parity/worker.log diff --git a/CHANGELOG.md b/CHANGELOG.md index b68e4d36..2f0ca98f 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -4,6 +4,51 @@ All notable changes to models are documented here. ## [Unreleased] +## [0.3.53] - 2026-10-07 + +### Added + +- `backend-cuda` is now published. It ships one `sm_80` PTX module serving every device of compute + capability 8.0 or above, carried in the jar under `META-INF/models/cuda/` with a SHA-256 the Java + loader recomputes. **It is opt-in and nothing activates it by accident**: there is no + `META-INF/services` entry, so no `ServiceLoader` discovers it, a consumer has to call + `CudaGgufBatchedMatrixKernel.open()` and inject the kernel, and `-Dmodels.cuda.disabled=true` is a + kill switch on top of that. Having the jar on the classpath does not route anything to a device. +- `CudaRoutingCounters.declined(type, stage, reason, rows, cols)`, with the shape in the key, and + `declinedProjections` / `totalDeclinedProjections` in the gate report. A projection that answered + no to `isEligible` and left through the Java branch previously incremented nothing, so an empty + refusals map read as "everything ran on the device" when it only meant "nothing hit an explicit + refusal path". + +### Fixed + +- **The FFN gate and up projections never reached the device.** `CudaGgufBatchedMatrixKernel` did + not override `multiplyDual`, so `isDualEligible` inherited the SPI default of `false` and + `LlamaForwardPass.dualMatmulDispatch` sent both to `TensorOps.ggufDualMatmul` on every layer of + every token. On Granite 4.1 3B they are 8192x2560 each -- 53.3% of a layer's projection + arithmetic. The triple path for query, key and value was implemented; the dual path was not, and + nothing recorded the asymmetry. + +### Measured + +Both device gates pass, on an RTX 4090 at compute capability 8.9, Granite 4.1 3B Q4_K_M, 20 prompts: + +| gate | result | evidence | +| --- | --- | --- | +| G1 token parity | **passed** -- 1280 token ids identical | `benchmark-results/2026-10-07-g1-parity` | +| G4 decode speed | **passed** -- 6.184x against a 3.00x gate | `benchmark-results/2026-10-07-g4-dualpath` | + +G4 moved from 1.788x to 6.184x, decode from 9.02 to 31.26 tok/s, with a control arm that moved +0.2% and a byte-identical PTX module either side. Routing the two projections raised dispatch -- +launches 241 to 321, transfers 402 to 522 per decode step -- and it did not matter. The earlier +reading of 1.788x as a dispatch ceiling was wrong; the binding constraint was the unimplemented +path. + +**One host and one model.** G1 has no tolerance, so each qualifying hardware profile and each +architecture family earns its own run before the claim generalises. No published Java artifact +behaviour changed in this release: `v0.3.52..HEAD` touches no other published module. + + ## [0.3.52] - 2026-10-06 ### Fixed diff --git a/RELEASING.md b/RELEASING.md index c87f3425..b92a0b50 100644 --- a/RELEASING.md +++ b/RELEASING.md @@ -6,11 +6,24 @@ the GitHub release. The publication allowlist contains `models-api`, `models-runtime`, `models`, `models-rag`, `models-semantic-order`, `backend-java`, `backend-tornado`, `backend-native`, -`backend-apple`, `models-langchain4j`, `models-spring-ai`, +`backend-cuda`, `backend-apple`, `models-langchain4j`, `models-spring-ai`, `models-spring-boot-starter`, `models-embedding`, `models-audio`, `models-router`, and `models-decisions`. Benchmark applications, documentation tooling, and modules containing only package scaffolding are not published. +`backend-cuda` publishes an **opt-in** artifact and nothing activates it by accident. There is no +`META-INF/services` entry, so no `ServiceLoader` discovers it: a consumer has to call +`CudaGgufBatchedMatrixKernel.open()` and inject the kernel, and `-Dmodels.cuda.disabled=true` is a +kill switch on top of that. One `sm_80` PTX module serves every device of compute capability 8.0 or +above, carried in the jar under `META-INF/models/cuda/` with a SHA-256 the Java loader recomputes. + +**Its numeric state must be stated in the release notes, not assumed from its presence.** G4 (decode +speed) passes at 6.184x against a 3.00x gate, measured in +`benchmark-results/2026-10-07-g4-dualpath`. G1 (exact token parity against the CPU path) is a +separate gate with no tolerance, and a release note must say which way it last ran and on what +hardware. Shipping the jar does not mean the device path is numerically verified; it means a caller +can opt into it and read the gates. + ## Cut a release 1. Set a non-snapshot version in `gradle.properties`. diff --git a/backend-cuda/build.gradle.kts b/backend-cuda/build.gradle.kts index d2daddfd..6318ce56 100644 --- a/backend-cuda/build.gradle.kts +++ b/backend-cuda/build.gradle.kts @@ -241,3 +241,32 @@ tasks.withType().configureEach { tasks.named("check") { dependsOn(cargoTestHost, cargoClippy, verifyPtxArtifact) } + +// Coverage: the 0.80 bar published modules carry applies to everything here that a host can +// execute, and two classes are exempted because they structurally cannot be. +// +// CudaDriver is the FFM binding to libcuda.so.1 -- every method is a downcall, so without a driver +// there is nothing to cover. CudaGgufBatchedMatrixKernel's bulk is the dispatch path behind those +// downcalls. Together they are 2,296 of the module's 2,491 missed instructions; the rest of the +// module measures 0.86 covered without them, and the classes a host *can* reach are already well +// past the bar -- CudaRoutingCounters at 391 of 403 instructions, Q8KActivations at 177 of 193. +// +// Their verification is hardware, not more off-device tests, and it is retained as evidence rather +// than asserted: G1 token parity and G4 decode speed in benchmark-results/2026-10-07-g1-parity and +// benchmark-results/2026-10-07-g4-dualpath, plus CudaQ6KDeviceParityTest, which launches the Q6_K +// kernel against the CPU control at six widths and skips off-device. Lowering the global bar to let +// this module in would have hidden a real gap in every other published module instead. +tasks.named("jacocoTestCoverageVerification") { + classDirectories.setFrom( + files( + classDirectories.files.map { directory -> + fileTree(directory) { + exclude( + "**/CudaDriver*.class", + "**/CudaGgufBatchedMatrixKernel*.class", + ) + } + }, + ), + ) +} diff --git a/benchmark-results/2026-10-07-g1-parity/README.md b/benchmark-results/2026-10-07-g1-parity/README.md new file mode 100644 index 00000000..83696ba1 --- /dev/null +++ b/benchmark-results/2026-10-07-g1-parity/README.md @@ -0,0 +1,60 @@ +# G1 passes: 1280 token ids identical, device against CPU + +**Measured 2026-10-07 on an RTX 4090 (compute capability 8.9, driver 13020), Runpod.** Models +revision `15661fb2836967e200322e5b0ceaa91416790994`. Kernel sha256 `4c534e622452df95...`, PTX target +`sm_80`. Model: Granite 4.1 3B Q4_K_M, sha256 recomputed by the gate from the file it opened. + +``` +PASS cuda-kernel-gate mode=parity device=NVIDIA GeForce RTX 4090 prompts=20 tokens=1280 identical +``` + +| gate | result | +| --- | --- | +| `tokenParity` (G1) | **passed** -- 1280 token ids identical across 20 prompts | +| `routingObservability` (G2) | passed -- 2,405,576 accelerated operations | +| `startupHonesty` (G5) | passed -- 344 ms readiness against a 120,000 ms ceiling | +| `qualified` | **true** | + +`parity.firstDivergence` is **null**: every token id sequence matched. `parity.selfTest` is +**false**, which is the field that makes this G1 evidence at all -- off-device both arms are the +Vector API and the command says `SELF-TEST` instead of `PASS`. One sequence hit end-of-generation, +which is recorded and deliberately not obeyed, so the compared length is not shortened where the +arms agree. + +`routing.totalDeclinedProjections` is **0**, from the counter added the same day. Nothing fell back +to the Java path, so the 1280 identical tokens were produced with the projections actually on the +device rather than by a quiet fallback that would have made parity trivial. + +## What changed since G1 last failed + +A note from a 2026-09-26 A40 run recorded G1 failing on an argmax flip at token 7, attributed to +`models_gqa_decode_attention`'s `expf`. Between then and now `expf` was rewritten on a different +principle, stated in `backend-cuda/src/main/rust/models-cuda-kernels/src/attention.rs`: + +> *"The requirement is not that it be accurate -- it is that it be **the same function the CPU +> runs**."* + +It transcribes the CPU's clamp, magic-constant rounding, Cody-Waite split, Taylor coefficients and +single-step exponent construction, in that order, with `fma` exactly where the CPU has one. That is +what exact token parity needed, and accuracy alone would not have given it. + +Two differences from the failing run are recorded rather than waved at: this is an RTX 4090 at +compute capability 8.9 where that was an A40 at 8.6, and this revision routes the FFN gate and up +projections to the device where that one left them on the CPU. The same `sm_80` module serves both +capabilities, but a single host does not establish hardware independence. **Re-run on an A40 or +L40S before claiming G1 holds across the qualifying profiles.** + +## Both gates now pass + +- **G1**, here: 1280/1280 token ids identical. +- **G4**, `benchmark-results/2026-10-07-g4-dualpath`: 6.184x decode against a 3.00x gate, 31.26 + against 5.06 tok/s, with a control arm that moved 0.2%. + +Numerically exact and 6.18x, on one host, on one model. + +## What this does not establish + +One model and one device. Granite 4.1 3B is 40 blocks of `granite` architecture with Q4_K and Q6_K +projections; it exercises neither a mixture-of-experts routed FFN nor a per-layer feed-forward width +nor an architecture whose attention scale differs. G1 is a no-tolerance gate, so each qualifying +hardware profile and each architecture family earns its own run. diff --git a/benchmark-results/2026-10-07-g1-parity/rust-ptx-parity.json b/benchmark-results/2026-10-07-g1-parity/rust-ptx-parity.json new file mode 100644 index 00000000..16259846 --- /dev/null +++ b/benchmark-results/2026-10-07-g1-parity/rust-ptx-parity.json @@ -0,0 +1,119 @@ +{ + "schemaVersion" : 2, + "createdAt" : "2026-10-07T19:29:53.432864217Z", + "policyVersion" : "gpu-large-model-rust-ptx-v2", + "modelsRevision" : "15661fb2836967e200322e5b0ceaa91416790994", + "mode" : "parity", + "accelerated" : true, + "refusalReason" : "", + "device" : { + "name" : "NVIDIA GeForce RTX 4090", + "computeCapability" : 89, + "globalMemoryBytes" : 25252724736, + "driverVersion" : 13020 + }, + "kernel" : { + "sha256" : "4c534e622452df95ca7fbb3accc65246be54e2c74d53714d98a03be6d31e8e0e", + "ptxTarget" : "sm_80", + "toolchain" : "nightly-2026-09-17", + "entryPoints" : [ "models_q4k_decode_projection", "models_q6k_decode_projection", "models_gqa_decode_attention" ] + }, + "environment" : { + "host" : "9342571901f5", + "osName" : "Linux", + "osVersion" : "6.8.0-146-generic", + "architecture" : "amd64", + "cpuModel" : "AMD EPYC 7452 32-Core Processor", + "processors" : 16, + "physicalMemoryBytes" : 142999998464, + "maxHeapBytes" : 32178700288, + "javaVersion" : "25.0.4.1", + "javaVendor" : "Eclipse Adoptium", + "vmName" : "OpenJDK 64-Bit Server VM" + }, + "configuration" : { + "mode" : "parity", + "modelPath" : "/work/model.gguf", + "modelSha256" : "662b0626cd58f443baea23559b469df6576a81d349649c59413b36a9fb32eb29", + "modelBytes" : 2099501664, + "promptsPath" : "/work/models/benchmark-results/2026-09-18-gpu-large-model/prompts.txt", + "promptsSha256" : "d0c0267f1eeb96b97a6e26f964d4e06c0e967ad7567e40881e54c791d74c7016", + "promptCount" : 20, + "maxTokens" : 64, + "warmupTokens" : 16, + "contextLength" : 4096, + "arm" : "both", + "sampling" : "greedy-argmax-lowest-index-v1", + "cudaDisabledProperty" : false, + "cudaDeviceOrdinal" : 0, + "jvmArguments" : [ "--add-modules=jdk.incubator.vector", "--enable-native-access=ALL-UNNAMED", "-XX:NativeMemoryTracking=summary" ], + "workingDirectory" : "/work/models" + }, + "routing" : { + "acceleratedOperations" : { + "F32/DECODE_ATTENTION" : 2040000, + "Q6_K/DECODE_PROJECTION" : 52296, + "Q4_K/PREFILL_PROJECTION" : 6240, + "Q6_K/PREFILL_PROJECTION" : 1040, + "Q4_K/DECODE_PROJECTION" : 306000 + }, + "refusals" : { }, + "declinedProjections" : { }, + "totalAcceleratedOperations" : 2405576, + "totalDeclinedProjections" : 0, + "inert" : false, + "kernelLaunches" : 416576, + "hostToDeviceTransfers" : 260456, + "deviceToHostTransfers" : 416576, + "hostToDeviceBytes" : 13756064960, + "deviceToHostBytes" : 8247738368, + "weightUploads" : 281, + "weightUploadBytes" : 2095104000, + "decodeSteps" : 1260, + "decodeProjections" : 358296, + "measuredDecodeSteps" : 1260, + "decodeStepsMarked" : true, + "launchesPerDecodeStep" : 321.0, + "transfersPerDecodeStep" : 522.0, + "activationBytesPerDecodeStep" : 1.5398952E7 + }, + "readinessMillis" : 344, + "parity" : { + "selfTest" : false, + "promptCount" : 20, + "tokensPerPrompt" : 64, + "comparedTokens" : 1280, + "sequencesHittingEndOfGeneration" : 1, + "identical" : true, + "firstDivergence" : null, + "promptDigests" : [ "bfe7bee92839632f1d68ebe8b65fb9cd3824ba3c926f8c6b0ccebec3571c985e", "9c98531297412ba6e3ebd3fc1bb3881552611d775b01b386d5c52258d96d4527", "decc00c635aa1608bc8affe58c034ea40ad5afea8712fcdcf4b350e1083d7b25", "6fae40b90f5a86c553ea914279f5c38de0fd8c6928d6e58b586c9af6e9e1c68d", "c6c78dc688a4443e094949c420f7e5e363ce737e52abf84e66f2aa2a20b83bc1", "a65c124623f65e639d05a06e5056d532e29532ae1952bdfeb408706f566834ab", "5f97c648fc5329865dac35d5fb20bb15667cd5a500c7752d60241db191aa7646", "3c61cb0871e38014fd1ffeab1e7c9b86fecae87dd39f8cc51daaeaf0ce3d5b1d", "85fa275abad2fcf7652db2060a0ca5ae85f2cb3ee895dbf9fdad3844277f215e", "c3d67901cd73e8dc1f6477e2da2316d0ceeb5c18d7a1ca2817c5e353936d9e20", "e3fb263583e2d93a9536dcd9bb735bb68e3d71c4e14ff704a7a5cea7abdb95ab", "55a4be086aef659ef78f2cf8d194c7aff228c95a2b23241a166b9f3842cb29e1", "0ba7006470ec8d2c1ec60056471ac68f000c8a3597c925b4497f76cc33c494f1", "531fe0c1a690b066dacf2ed30e63cd86737cb06615bb99a39f5d302d7f65d5e7", "16422f3a476e41cbe64164b421b9b66f97d07e82a22c240cff332ba328b81bd0", "00a434f63a5b3b89005ec8e98e7eb37ed2dc5fa6300533165b31df949314b102", "574cd13af5ceb84475eab456207f3cb0b258760a79a4228c631dabe5c0022388", "72a215d0a7efd7fae7dadc7b633701204d9f4bafe4487f302f5c13877aa9a444", "9fd82302b33caf3033c421ef429e69b273f487fac4f4ba2744905e26511c3a24", "9618c19f55a9736ee7cc5cac3f521d92762e3c1a28226e83194008b9b0c27ee4" ] + }, + "decode" : null, + "memory" : { + "peakProcessRssBytes" : 3707437056, + "residentBytes" : 1690697728, + "heapUsedBytes" : 377563128, + "peakDeviceBytes" : 2097012576, + "deviceTotalMemoryBytes" : 25252724736, + "reading" : "peakDeviceBytes covers kernel-owned allocations only (weights, staging scratch, per-call attention buffers); the CUDA context and loaded module are not included, so it is a lower bound. peakProcessRssBytes is process-wide and therefore not separable per arm in a two-arm run: use --arm accelerated and --arm control in separate processes for per-arm host peaks." + }, + "gates" : { + "routingObservability" : { + "gate" : "G2", + "passed" : true, + "evidence" : "2405576 accelerated operations" + }, + "startupHonesty" : { + "gate" : "G5", + "passed" : true, + "evidence" : "344 ms readiness against a 120000 ms ceiling" + }, + "tokenParity" : { + "gate" : "G1", + "passed" : true, + "evidence" : "1280 token ids identical across 20 prompts" + }, + "decodeSpeedup" : null + }, + "qualified" : true +} diff --git a/benchmark-results/2026-10-07-g1-parity/worker.log b/benchmark-results/2026-10-07-g1-parity/worker.log new file mode 100644 index 00000000..b8471508 --- /dev/null +++ b/benchmark-results/2026-10-07-g1-parity/worker.log @@ -0,0 +1,71 @@ +[19:21:04] nvidia-smi: +NVIDIA GeForce RTX 4090, 8.9, 595.91.07 +[19:21:04] installing jdk 25 +[19:21:09] jdk: openjdk version "25.0.4.1" 2026-08-18 LTS +[19:21:09] installing pinned rust nightly-2026-09-17 +[19:21:30] rust: rustc 1.100.0-nightly (923c95cdf 2026-09-16) +[19:21:30] fetching source +[19:21:33] source at /work/models, revision 15661fb2836967e200322e5b0ceaa91416790994 +[19:21:33] === STEP 1: ptxas assembly (cheap gate) === + Downloaded getopts v0.2.24 + Downloaded rustc-demangle v0.1.28 + Downloaded hashbrown v0.17.1 + Downloaded libc v0.2.189 + Compiling compiler_builtins v0.1.160 (/root/.rustup/toolchains/nightly-2026-09-17-x86_64-unknown-linux-gnu/lib/rustlib/src/rust/library/compiler-builtins/compiler-builtins) + Compiling core v0.0.0 (/root/.rustup/toolchains/nightly-2026-09-17-x86_64-unknown-linux-gnu/lib/rustlib/src/rust/library/core) + Compiling models-cuda-kernels v0.3.46 (/work/models/backend-cuda/src/main/rust/models-cuda-kernels) + Finished `release` profile [optimized] target(s) in 21.85s +[19:23:48] ptx: backend-cuda/build/generated/cuda-resources/META-INF/models/cuda/models-cuda-kernels.ptx (64104 bytes) +[19:23:48] ptxas OK +[19:23:48] === STEP 2: capability gate === +WARNING: Using incubator modules: jdk.incubator.vector +PASS cuda-kernel-gate mode=capability device=NVIDIA GeForce RTX 4090 cc=89 kernel=4c534e622452 readiness=527 ms report=/work/capability.json +[19:23:50] capability report uploaded +[19:23:50] === STEP 2a: Q6_K device parity (gate before any model) === +CudaQ6KDeviceParityTest > a Q6_K row projection is bit-exact against the CPU control at every width > 1 super-block(s) PASSED + +CudaQ6KDeviceParityTest > a Q6_K row projection is bit-exact against the CPU control at every width > 2 super-block(s) PASSED + +CudaQ6KDeviceParityTest > a Q6_K row projection is bit-exact against the CPU control at every width > 3 super-block(s) PASSED + +CudaQ6KDeviceParityTest > a Q6_K row projection is bit-exact against the CPU control at every width > 32 super-block(s) PASSED + +CudaQ6KDeviceParityTest > a Q6_K row projection is bit-exact against the CPU control at every width > 33 super-block(s) PASSED + +CudaQ6KDeviceParityTest > a Q6_K row projection is bit-exact against the CPU control at every width > 48 super-block(s) PASSED + +BUILD SUCCESSFUL in 11s +13 actionable tasks: 3 executed, 10 up-to-date +Consider enabling configuration cache to speed up this build: https://docs.gradle.org/9.4.1/userguide/configuration_cache_enabling.html +[19:24:01] Q6_K device parity OK +[19:24:01] === fetching model (granite 4.1 3b q4_k_m, 1.96GB) === +[19:24:27] model sha verified +[19:24:27] === STEP 4: parity gate G1 === +WARNING: Using incubator modules: jdk.incubator.vector +Oct 07, 2026 7:24:30 PM com.integrallis.vectors.core.PanamaVectorUtilSupportProvider create +INFO: vectors-core: Using Panama Vector API SIMD provider (vector bits: 256) +Oct 07, 2026 7:24:30 PM com.integrallis.vectors.core.VectorizationProvider logActiveToggles +INFO: vectors-core: provider=PanamaVectorUtilSupport panama=true maxBits=256 preferredBits=256 fastVectorFMA=true fastScalarFMA=true sve=false ggufParallel=true ggufParallelThreshold=1048576 ggufExecutor=persistent ggufThreads=16 ggufChunksPerThread=2 mappedKQuantLongOffsets=auto(q4=true,q5=false,q6=false) q4ShortPairwiseSupported=true q4UnsignedPairwiseSupported=true toggles=[(defaults ? no -Dvectors.* overrides)] +PASS cuda-kernel-gate mode=parity device=NVIDIA GeForce RTX 4090 prompts=20 tokens=1280 identical report=/work/parity.json +[19:29:54] decode report uploaded +[19:29:54] parity gate rc=0 +[19:29:54] === the numbers this run exists for === + --- G1 verdict --- + accelerated = True + refusalReason = + selfTest = False + qualified = True + tokenParity: passed=True | 1280 token ids identical across 20 prompts + --- sequences --- + parity.promptCount = 20 + parity.sequencesHittingEndOfGeneration = 1 + --- first divergence --- + none: every token id sequence identical + --- routing --- + routing.kernelLaunches = 416576 + routing.decodeProjections = 358296 + routing.totalDeclinedProjections = 0 + routing.launchesPerDecodeStep = 321.0 + routing.transfersPerDecodeStep = 522.0 + declined: none +[19:29:54] FINAL_STATUS=COMPLETE diff --git a/build.gradle.kts b/build.gradle.kts index e9412237..d3ddbcb2 100644 --- a/build.gradle.kts +++ b/build.gradle.kts @@ -53,6 +53,7 @@ val publishedModuleNames = "backend-java", "backend-tornado", "backend-native", + "backend-cuda", "backend-apple", "models-langchain4j", "models-spring-ai", diff --git a/gradle.properties b/gradle.properties index 8e4b60b5..40869a9d 100644 --- a/gradle.properties +++ b/gradle.properties @@ -1,2 +1,2 @@ -version = 0.3.52 +version = 0.3.53 vectorsVersion = 0.1.28 From 10af16c33be7e023c114812e6137b47fab7e0c90 Mon Sep 17 00:00:00 2001 From: Brian Sam-Bodden Date: Wed, 7 Oct 2026 12:38:18 -0700 Subject: [PATCH 2/5] Ignore .claude/, which git add -A once swept into a release commit The agent worktrees under .claude/worktrees are embedded git repositories, so adding them produces a commit that clones cannot resolve. Caught during 0.3.53 preparation and removed from that commit; ignored here so it cannot recur. --- .gitignore | 4 ++++ 1 file changed, 4 insertions(+) diff --git a/.gitignore b/.gitignore index 6853079b..d2bc9611 100644 --- a/.gitignore +++ b/.gitignore @@ -50,3 +50,7 @@ docs/content/modules/ROOT/attachments/javadoc/ # Model cache directory **/.jvllm/ + +# Agent worktrees and session scratch. These are embedded git repositories; git add -A +# swept one into a release commit once. +.claude/ From 7983a333c01e3a800c3b63f2033f80fa57042ecd Mon Sep 17 00:00:00 2001 From: Brian Sam-Bodden Date: Wed, 7 Oct 2026 13:36:43 -0700 Subject: [PATCH 3/5] Retain the backend-tornado hardware run: exit 0, and 36 failed device launches RELEASING.md requires the public loader and exact CPU/GPU parity gate on each qualified NVIDIA profile, retained under models-accelerator-bench/results/. This is that run on an RTX 4090 with TornadoVM v5.2.0-jdk25 built from source for the PTX backend, and it is retained as a failed precondition rather than a pass. The gate exited 0 and reported accelerated=true device=cuda-0 reason=eligible, readiness 32,171 ms, plans=321, prefill 69.34 tok/s, decode 6.49 tok/s, and routed projections {Q4_K=16800, Q6_K=2870}. Both format counts being non-zero is the check the models-bench README asks for, so the mixed-format partial routing trap is not what happened. What the report does not contain is 36 occurrences of [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 CUDA 701 is CUDA_ERROR_LAUNCH_OUT_OF_RESOURCES. Thirty-six launches failed on the device, and the JSON has no field for a failed launch, so a reader of the artifact alone sees accelerated=true and 6.49 tok/s. Exit status is therefore not evidence the device path worked, and the throughput is not filed as a measurement of accelerated decode. The number supports that reading. 6.49 tok/s is barely above the 5.06 tok/s Vector API control measured on the same model and GPU class in benchmark-results/2026-10-07-g4-dualpath, where backend-cuda reached 31.26. Readiness of 32,171 ms is inside G5's 120,000 ms ceiling and two orders of magnitude above backend-cuda's 344-353 ms on the same hardware. This is the third instance in one day of evidence reporting success while something silently did not run: backend-cuda left 53.3% of a layer's projection arithmetic on the CPU behind an empty refusals map; the counter added to catch that printed None because it was never serialised; and here a gate returns 0 over 36 failed launches. The fix is the same shape each time -- the artifact has to carry the negative. accelerator-profile needs a launch-failure count, and --require true should fail on a nonzero one. Also recorded for whoever runs this next, because each cost a pod: the pinned tag builds with --jdk jdk25, which yields -Pjdk25,ptx-backend, and the wrong value compiles at -source 8 and dies on sealed classes; the assembled SDK lands under dist/tornadovm---/tornadovm--/ and TORNADOVM_HOME must point there because bin/tornado reads $TORNADOVM_HOME/etc/tornado.backend; and :models-bench:installDist pulls in :backend-cuda:compilePtx, so cargo and the pinned nightly's rust-src are required. --- ...90-ptx-2026-10-07-accelerator-profile.json | 20 +++ .../results/rtx4090-ptx-2026-10-07-worker.log | 121 ++++++++++++++++++ .../results/rtx4090-ptx-2026-10-07.md | 83 ++++++++++++ 3 files changed, 224 insertions(+) create mode 100644 models-accelerator-bench/results/rtx4090-ptx-2026-10-07-accelerator-profile.json create mode 100644 models-accelerator-bench/results/rtx4090-ptx-2026-10-07-worker.log create mode 100644 models-accelerator-bench/results/rtx4090-ptx-2026-10-07.md diff --git a/models-accelerator-bench/results/rtx4090-ptx-2026-10-07-accelerator-profile.json b/models-accelerator-bench/results/rtx4090-ptx-2026-10-07-accelerator-profile.json new file mode 100644 index 00000000..4743db3a --- /dev/null +++ b/models-accelerator-bench/results/rtx4090-ptx-2026-10-07-accelerator-profile.json @@ -0,0 +1,20 @@ +{ + "timestamp" : "2026-10-07T20:33:51.220769153Z", + "model" : "/work/model.gguf", + "accelerated" : true, + "device" : "cuda-0", + "reason" : "eligible", + "requiredBytes" : 4467438784, + "readinessMillis" : 32171, + "projectionPlans" : 321, + "routedProjectionsByFormat" : { + "Q4_K" : 16800, + "Q6_K" : 2870 + }, + "promptTokens" : 26, + "generatedTokens" : 64, + "prefillTokensPerSecond" : 69.3391006823717, + "decodeTokensPerSecond" : 6.492315709531025, + "prefillMillis" : 374.968809, + "decodeMillis" : 9857.807732 +} \ No newline at end of file diff --git a/models-accelerator-bench/results/rtx4090-ptx-2026-10-07-worker.log b/models-accelerator-bench/results/rtx4090-ptx-2026-10-07-worker.log new file mode 100644 index 00000000..fc800b18 --- /dev/null +++ b/models-accelerator-bench/results/rtx4090-ptx-2026-10-07-worker.log @@ -0,0 +1,121 @@ +[20:22:39] nvidia-smi: +NVIDIA GeForce RTX 4090, 8.9, 570.158.01 +Cuda compilation tools, release 12.8, V12.8.93 +Build cuda_12.8.r12.8/compiler.35583870_0 +[20:22:39] installing build prerequisites +[20:26:06] maven: Apache Maven 3.8.7 +[20:26:06] cmake: cmake version 3.28.3 +[20:26:06] installing the pinned rust nightly -- :models-bench:installDist pulls in +[20:26:06] :backend-cuda:compilePtx, which needs cargo and -Zbuild-std's rust-src +[20:26:24] cargo: cargo 1.100.0-nightly (495c385d0 2026-09-16) +[20:26:24] installing jdk 25 +[20:26:28] jdk: openjdk version "25.0.4.1" 2026-08-18 LTS +[20:26:28] === building TornadoVM v5.2.0-jdk25 from source, PTX backend === +[20:26:28] TORNADOVM_HOME=/work/tornado-sdk (an input to bin/compile, not an output) +[20:26:30] tornado tree at v5.2.0-jdk25 +[20:26:30] building via bin/compile --jdk jdk25 --backend ptx (bin/compile builds -P,; the wrong jdk value gives -source 8) +[20:27:58] tornado build rc=0 +20:27:58 [INFO] tornado-cufft ...................................... SUCCESS [ 1.567 s] +20:27:58 [INFO] tornado-cudnn ...................................... SUCCESS [ 1.570 s] +20:27:58 [INFO] tornado-cusparse ................................... SUCCESS [ 1.151 s] +20:27:58 [INFO] tornado-cutlass .................................... SUCCESS [ 1.260 s] +20:27:58 [INFO] tornado-unittests .................................. SUCCESS [ 5.187 s] +20:27:58 [INFO] tornado-annotation ................................. SUCCESS [ 1.372 s] +20:27:58 [INFO] tornado-assembly ................................... SUCCESS [ 28.687 s] +20:27:58 [INFO] ------------------------------------------------------------------------ +20:27:58 [INFO] BUILD SUCCESS +20:27:58 [INFO] ------------------------------------------------------------------------ +20:27:58 [INFO] Total time: 01:05 min (Wall Clock) +20:27:58 [INFO] Finished at: 2026-10-07T20:27:58Z +20:27:58 [INFO] ------------------------------------------------------------------------ +Maven build succeeded +########################################################################### +TornadoVM build success +Updating PATH and TORNADOVM_HOME to tornadovm-5.2.0-jdk25-ptx-linux-amd64 +Backend : PTX +Commit : db146ae +########################################################################### +Generated utilities: + [Unix-env]: /work/tornadovm/setvars.sh + [argfile]: /work/tornadovm/dist/tornadovm-5.2.0-jdk25-ptx-linux-amd64/tornadovm-5.2.0-jdk25-ptx/tornado-argfile +[INFO] Run first: source setvars.sh + +[20:27:58] not at $TORNADOVM_HOME/bin/tornado, searching the assembled dist +[20:27:58] launcher=/work/tornadovm/dist/tornadovm-5.2.0-jdk25-ptx-linux-amd64/tornadovm-5.2.0-jdk25-ptx/bin/tornado +[20:27:58] TORNADOVM_HOME now /work/tornadovm/dist/tornadovm-5.2.0-jdk25-ptx-linux-amd64/tornadovm-5.2.0-jdk25-ptx +[20:27:58] backends: tornado.backends=ptx-backend +version=5.2.0-jdk25 +branch=UNKNOWN +commit=db146ae +WARNING: Using incubator modules: jdk.incubator.vector + +Number of Tornado drivers: 1 +Driver: PTX + Total number of PTX devices : 1 + Tornado device=0:0 (DEFAULT) + PTX -- PTX -- NVIDIA GeForce RTX 4090 + Global Memory Size: 23.5 GB + Local Memory Size: 48.0 KB + Workgroup Dimensions: 3 + Total Number of Block Threads: [2147483647, 65535, 65535] + Max WorkGroup Configuration: [1024, 1024, 64] + Device OpenCL C version: N/A + + +[20:27:59] === fetching models source === +[20:28:03] models at 10af16c33be7e023c114812e6137b47fab7e0c90 +[20:28:03] === building models-bench dist === +[20:29:54] installDist rc=0 +[20:29:54] === fetching model === +[20:33:06] model sha verified +[20:33:06] === accelerator-profile, the documented gate command === +[20:33:52] accelerator-profile rc=0 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 +accelerator profile: accelerated=true device=cuda-0 reason=eligible readiness=32171 ms plans=321 + prompt=26 tokens prefill=69.34 tok/s (375.0 ms) decode=64 tokens 6.49 tok/s (9857.8 ms) + routed projections by format: {Q4_K=16800, Q6_K=2870} +report: /work/accelerator-profile.json +[20:33:53] profile uploaded +[20:33:53] === per-format routing, the point of the command === + accelerated = True reason = eligible + device = cuda-0 + readiness = 32171 + prefillTokensPerSecond = 69.3391006823717 + decodeTokensPerSecond = 6.492315709531025 + routedProjectionsByFormat = {"Q4_K": 16800, "Q6_K": 2870} +[20:33:53] FINAL_STATUS=COMPLETE diff --git a/models-accelerator-bench/results/rtx4090-ptx-2026-10-07.md b/models-accelerator-bench/results/rtx4090-ptx-2026-10-07.md new file mode 100644 index 00000000..38812845 --- /dev/null +++ b/models-accelerator-bench/results/rtx4090-ptx-2026-10-07.md @@ -0,0 +1,83 @@ +# backend-tornado on RTX 4090 / PTX: the gate returns 0 and the device errors 36 times + +**Measured 2026-10-07.** RTX 4090, compute capability 8.9. TornadoVM **v5.2.0-jdk25** built from +source with the PTX backend, JDK 25. Models revision `10af16c33be7e023c114812e6137b47fab7e0c90`. +Model: Granite 4.1 3B Q4_K_M, sha256 verified on download. The `RELEASING.md` precondition names +this directory, so the run is retained here either way. + +Command, exactly as `models-bench/README.md` documents it: + +``` +tornado -cp "$classpath" \ + --params="accelerator-profile --model /work/model.gguf --tokens 64 --batch 32 --require true \ + --output /work/accelerator-profile.json" \ + com.integrallis.models.bench.InferenceBenchmarkCli +``` + +## What it reported + +``` +accelerator profile: accelerated=true device=cuda-0 reason=eligible readiness=32171 ms plans=321 + prompt=26 tokens prefill=69.34 tok/s (375.0 ms) decode=64 tokens 6.49 tok/s (9857.8 ms) + routed projections by format: {Q4_K=16800, Q6_K=2870} +``` + +Exit status 0. `--require true` was set, so an ineligible accelerator would have been a hard +failure; it was not reached. + +**Both format counts are non-zero**, which is the check the README asks for: *"a run whose `Q6_K` +count is zero routed only part of the model while still producing a plausible throughput number."* +Q4_K 16,800 and Q6_K 2,870, so the mixed-format trap is not what happened here. + +## What it does not report, and why this is not a clean pass + +The run emitted **36 occurrences** of + +``` +[TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 +``` + +CUDA 701 is `CUDA_ERROR_LAUNCH_OUT_OF_RESOURCES`. Thirty-six kernel launches failed on the device. + +**The report JSON records none of them.** Its keys are `accelerated`, `decodeMillis`, +`decodeTokensPerSecond`, `device`, `generatedTokens`, `model`, `prefillMillis`, +`prefillTokensPerSecond`, `projectionPlans`, `promptTokens`, `readinessMillis`, `reason`, +`requiredBytes`, `routedProjectionsByFormat`, `timestamp`. There is no field for a failed launch, so +`accelerated: true` and `decodeTokensPerSecond: 6.49` are what a reader of the artifact alone would +take away. + +So the gate's exit status is not evidence that the device path worked. **Treat this as a failed +precondition with a known cause, not as a pass**, and do not file the throughput figure as a +measurement of accelerated decode. + +The throughput supports that reading rather than contradicting it. Decode at **6.49 tok/s** is +barely above the **5.06 tok/s** Vector API control measured on the same model and the same GPU class +in `benchmark-results/2026-10-07-g4-dualpath`, where the `backend-cuda` arm reached **31.26 tok/s**. +A device arm that errors 36 times and lands within 30% of its own CPU control is behaving like one +whose launches mostly failed. + +Readiness was **32,171 ms** against G5's 120,000 ms ceiling -- inside budget, and two orders of +magnitude above `backend-cuda`'s 344-353 ms on the same hardware. + +## The pattern, which is the transferable part + +This is the third instance in one day of evidence reporting success while something silently did not +run: + +1. `backend-cuda` left the FFN gate and up projections on the CPU -- 53.3% of a layer's projection + arithmetic -- because `multiplyDual` was unimplemented, and an empty `refusals` map read as + "nothing fell back". +2. The counter added to catch that printed `None`, because it was never serialised into the report. +3. Here, 36 failed `cuLaunchKernel` calls with `accelerated: true` and no field to hold them. + +The fix in each case is the same shape: the artifact has to carry the negative. `accelerator-profile` +needs a launch-failure count, and `--require true` should fail on a nonzero one. + +## Next, in order + +1. Give `accelerator-profile` a launch-failure counter and make `--require true` respect it, so this + cannot read as a pass again. +2. Diagnose the 701 itself. `plans=321` retained execution plans and a 32-second readiness suggest + per-plan device resources rather than model size; `--batch 32` and the retained-plan count are the + first two things to vary. +3. Re-run only after both, and only then record a `backend-tornado` verdict. From d68509d9d4719c6d3587a98d00a465a4bdfa90dd Mon Sep 17 00:00:00 2001 From: Brian Sam-Bodden Date: Wed, 7 Oct 2026 14:48:07 -0700 Subject: [PATCH 4/5] Remove backend-tornado: Models ships one GPU implementation, the better one Measured on the same RTX 4090 and the same model on the same day: backend-cuda backend-tornado G1 exact parity 1280/1280 not reached decode 31.26 tok/s 6.49 tok/s CPU control 5.06 tok/s 5.06 tok/s readiness 344 ms 32,171 ms device launch errors 0 36 x CUDA 701 self-contained yes no A second implementation of the same capability, five times slower than the first and barely ahead of the CPU path it exists to accelerate, is not worth the surface it costs. The 36 cuLaunchKernel failures are CUDA_ERROR_LAUNCH_OUT_OF_RESOURCES and its report had no field to record them, so the gate returned 0 with accelerated=true over a device that was erroring. It could not be self-contained either: a Maven artifact cannot carry TornadoVM's device runtime, so an application had to install a matching distribution and launch through its own launcher, which defeats the point of adding a dependency. Removed with the module, because each existed only to serve it: the accelerator-profile command and its test, the TornadoVM benchmark arm in models-accelerator-bench with its orphaned tests, the tornado-api and tornado-runtime dependencies, and the RELEASING.md precondition requiring a loader and parity gate on each NVIDIA profile under models-accelerator-bench/results/. That precondition existed for backend-tornado; backend-cuda's gates are cuda-kernel-gate --mode capability|parity|decode and need no external runtime. Docs are rewritten rather than left stale: README, modules.adoc, index.adoc, architecture.adoc, models-accelerator-bench/README and gpu-acceleration.adoc, which was 188 lines of TornadoVM and is now 107 documenting backend-cuda with its measured gates. Deliberately kept: the KV-ridge experiment that happened to live in models-accelerator-bench, which is unrelated work, and the August measurement records that name backend-tornado, which are what was measured then and are not rewritten to match a later decision. The 0.3.53 notes said no published module's behaviour changed. That was true when written and is not now, so it is corrected: the published surface changes in exactly two ways, backend-cuda added and backend-tornado removed. clean spotlessCheck build: 466 classes, 2587 tests, zero failures. --- CHANGELOG.md | 28 +- README.md | 19 +- RELEASING.md | 12 +- backend-tornado/.github/badges/mfcqi.json | 7 - backend-tornado/README.md | 35 - backend-tornado/build.gradle.kts | 55 - backend-tornado/gradle.lockfile | 75 -- .../tornado/AcceleratorEligibility.java | 320 ----- .../models/backend/tornado/DeviceBudget.java | 101 -- .../backend/tornado/DeviceMemoryRequest.java | 179 --- .../tornado/KQuantProjectionKernel.java | 568 --------- .../backend/tornado/PlanShapeStrategy.java | 108 -- .../backend/tornado/Q4ProjectionKernel.java | 348 ------ .../backend/tornado/TornadoBackend.java | 172 --- .../tornado/TornadoBackendOptions.java | 35 - .../tornado/TornadoBackendRuntime.java | 81 -- .../backend/tornado/TornadoBackendStatus.java | 48 - .../TornadoGgufBatchedMatrixKernel.java | 1091 ----------------- .../TornadoPureJavaBackendProvider.java | 72 -- .../tornado/TornadoRuntimeDevices.java | 52 - ...ckend.purejava.spi.PureJavaBackendProvider | 1 - .../tornado/AcceleratorEligibilityTest.java | 80 -- .../tornado/KQuantProjectionKernelTest.java | 697 ----------- .../tornado/LargeModelEligibilityTest.java | 292 ----- .../tornado/Q4ProjectionKernelTest.java | 200 --- .../TornadoBackendIntegrationTest.java | 59 - .../tornado/TornadoBackendOptionsTest.java | 40 - .../tornado/TornadoBackendStatusTest.java | 55 - .../backend/tornado/TornadoBackendTest.java | 50 - .../TornadoGgufBatchedMatrixKernelTest.java | 210 ---- .../2026-10-07-tornado-removal/README.md | 25 +- .../accelerator-profile.json | 0 .../2026-10-07-tornado-removal/worker.log | 0 build.gradle.kts | 1 - .../modules/ROOT/pages/architecture.adoc | 10 +- .../modules/ROOT/pages/gpu-acceleration.adoc | 245 ++-- docs/content/modules/ROOT/pages/index.adoc | 14 +- docs/content/modules/ROOT/pages/modules.adoc | 9 +- models-accelerator-bench/README.md | 7 +- models-accelerator-bench/build.gradle.kts | 3 - .../CausalAttentionExperiment.java | 198 --- .../accelerator/CausalAttentionKernel.java | 238 ---- .../DeviceInventoryExperiment.java | 68 - .../LargeModelBudgetExperiment.java | 212 ---- .../Q4GroupedProjectionExperiment.java | 195 --- .../accelerator/Q4ProjectionExperiment.java | 330 ----- .../accelerator/Q8ProjectionExperiment.java | 241 ---- .../accelerator/Q8ProjectionKernel.java | 73 -- .../accelerator/QwenFullModelExperiment.java | 236 ---- .../accelerator/TornadoBackendExperiment.java | 107 -- .../TornadoBatchedCausalAttentionKernel.java | 178 --- .../TornadoCausalAttentionPlan.java | 197 --- .../CausalAttentionKernelTest.java | 181 --- .../accelerator/Q8ProjectionKernelTest.java | 100 -- .../QwenFullModelExperimentTest.java | 37 - models-bench/README.md | 41 +- models-bench/build.gradle.kts | 5 - .../models/bench/AcceleratorProfileCli.java | 265 ---- .../models/bench/InferenceBenchmarkCli.java | 4 - .../bench/AcceleratorProfileCliTest.java | 101 -- settings.gradle.kts | 1 - 61 files changed, 182 insertions(+), 8230 deletions(-) delete mode 100644 backend-tornado/.github/badges/mfcqi.json delete mode 100644 backend-tornado/README.md delete mode 100644 backend-tornado/build.gradle.kts delete mode 100644 backend-tornado/gradle.lockfile delete mode 100644 backend-tornado/src/main/java/com/integrallis/models/backend/tornado/AcceleratorEligibility.java delete mode 100644 backend-tornado/src/main/java/com/integrallis/models/backend/tornado/DeviceBudget.java delete mode 100644 backend-tornado/src/main/java/com/integrallis/models/backend/tornado/DeviceMemoryRequest.java delete mode 100644 backend-tornado/src/main/java/com/integrallis/models/backend/tornado/KQuantProjectionKernel.java delete mode 100644 backend-tornado/src/main/java/com/integrallis/models/backend/tornado/PlanShapeStrategy.java delete mode 100644 backend-tornado/src/main/java/com/integrallis/models/backend/tornado/Q4ProjectionKernel.java delete mode 100644 backend-tornado/src/main/java/com/integrallis/models/backend/tornado/TornadoBackend.java delete mode 100644 backend-tornado/src/main/java/com/integrallis/models/backend/tornado/TornadoBackendOptions.java delete mode 100644 backend-tornado/src/main/java/com/integrallis/models/backend/tornado/TornadoBackendRuntime.java delete mode 100644 backend-tornado/src/main/java/com/integrallis/models/backend/tornado/TornadoBackendStatus.java delete mode 100644 backend-tornado/src/main/java/com/integrallis/models/backend/tornado/TornadoGgufBatchedMatrixKernel.java delete mode 100644 backend-tornado/src/main/java/com/integrallis/models/backend/tornado/TornadoPureJavaBackendProvider.java delete mode 100644 backend-tornado/src/main/java/com/integrallis/models/backend/tornado/TornadoRuntimeDevices.java delete mode 100644 backend-tornado/src/main/resources/META-INF/services/com.integrallis.models.backend.purejava.spi.PureJavaBackendProvider delete mode 100644 backend-tornado/src/test/java/com/integrallis/models/backend/tornado/AcceleratorEligibilityTest.java delete mode 100644 backend-tornado/src/test/java/com/integrallis/models/backend/tornado/KQuantProjectionKernelTest.java delete mode 100644 backend-tornado/src/test/java/com/integrallis/models/backend/tornado/LargeModelEligibilityTest.java delete mode 100644 backend-tornado/src/test/java/com/integrallis/models/backend/tornado/Q4ProjectionKernelTest.java delete mode 100644 backend-tornado/src/test/java/com/integrallis/models/backend/tornado/TornadoBackendIntegrationTest.java delete mode 100644 backend-tornado/src/test/java/com/integrallis/models/backend/tornado/TornadoBackendOptionsTest.java delete mode 100644 backend-tornado/src/test/java/com/integrallis/models/backend/tornado/TornadoBackendStatusTest.java delete mode 100644 backend-tornado/src/test/java/com/integrallis/models/backend/tornado/TornadoBackendTest.java delete mode 100644 backend-tornado/src/test/java/com/integrallis/models/backend/tornado/TornadoGgufBatchedMatrixKernelTest.java rename models-accelerator-bench/results/rtx4090-ptx-2026-10-07.md => benchmark-results/2026-10-07-tornado-removal/README.md (77%) rename models-accelerator-bench/results/rtx4090-ptx-2026-10-07-accelerator-profile.json => benchmark-results/2026-10-07-tornado-removal/accelerator-profile.json (100%) rename models-accelerator-bench/results/rtx4090-ptx-2026-10-07-worker.log => benchmark-results/2026-10-07-tornado-removal/worker.log (100%) delete mode 100644 models-accelerator-bench/src/main/java/com/integrallis/models/accelerator/CausalAttentionExperiment.java delete mode 100644 models-accelerator-bench/src/main/java/com/integrallis/models/accelerator/CausalAttentionKernel.java delete mode 100644 models-accelerator-bench/src/main/java/com/integrallis/models/accelerator/DeviceInventoryExperiment.java delete mode 100644 models-accelerator-bench/src/main/java/com/integrallis/models/accelerator/LargeModelBudgetExperiment.java delete mode 100644 models-accelerator-bench/src/main/java/com/integrallis/models/accelerator/Q4GroupedProjectionExperiment.java delete mode 100644 models-accelerator-bench/src/main/java/com/integrallis/models/accelerator/Q4ProjectionExperiment.java delete mode 100644 models-accelerator-bench/src/main/java/com/integrallis/models/accelerator/Q8ProjectionExperiment.java delete mode 100644 models-accelerator-bench/src/main/java/com/integrallis/models/accelerator/Q8ProjectionKernel.java delete mode 100644 models-accelerator-bench/src/main/java/com/integrallis/models/accelerator/QwenFullModelExperiment.java delete mode 100644 models-accelerator-bench/src/main/java/com/integrallis/models/accelerator/TornadoBackendExperiment.java delete mode 100644 models-accelerator-bench/src/main/java/com/integrallis/models/accelerator/TornadoBatchedCausalAttentionKernel.java delete mode 100644 models-accelerator-bench/src/main/java/com/integrallis/models/accelerator/TornadoCausalAttentionPlan.java delete mode 100644 models-accelerator-bench/src/test/java/com/integrallis/models/accelerator/CausalAttentionKernelTest.java delete mode 100644 models-accelerator-bench/src/test/java/com/integrallis/models/accelerator/Q8ProjectionKernelTest.java delete mode 100644 models-accelerator-bench/src/test/java/com/integrallis/models/accelerator/QwenFullModelExperimentTest.java delete mode 100644 models-bench/src/main/java/com/integrallis/models/bench/AcceleratorProfileCli.java delete mode 100644 models-bench/src/test/java/com/integrallis/models/bench/AcceleratorProfileCliTest.java diff --git a/CHANGELOG.md b/CHANGELOG.md index 2f0ca98f..2f0d108d 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -6,6 +6,27 @@ All notable changes to models are documented here. ## [0.3.53] - 2026-10-07 +### Removed + +- **`backend-tornado` is removed.** Models ships one GPU implementation, and this was the weaker of + two. Measured on the same RTX 4090 and the same model on the same day: `backend-cuda` reached + exact token parity (1280 of 1280 token ids identical) and **31.26 tok/s** decode against a + **5.06 tok/s** Vector API control, with 344 ms readiness and no device errors. The TornadoVM arm + managed **6.49 tok/s** -- barely above the CPU path it exists to accelerate -- over **36 failed + `cuLaunchKernel` calls** (`CUDA_ERROR_LAUNCH_OUT_OF_RESOURCES`) that its report did not record, + with 32,171 ms readiness. It also could not be self-contained: the Maven artifact cannot carry + TornadoVM's device runtime, so an application had to install a matching distribution and launch + through its own launcher, which defeats "add a dependency and get acceleration". Evidence: + `benchmark-results/2026-10-07-tornado-removal`. +- Removed with it, because they existed only to serve it: the `accelerator-profile` command of + `models-bench`, the TornadoVM benchmark arm in `models-accelerator-bench`, the `tornado-api` and + `tornado-runtime` dependencies, and the release precondition in `RELEASING.md` that required a + loader and parity gate on each NVIDIA profile under `models-accelerator-bench/results/`. That + precondition existed for `backend-tornado`; `backend-cuda`'s gates are + `cuda-kernel-gate --mode capability|parity|decode` and need no external runtime. +- The KV-ridge experiment in `models-accelerator-bench` is unaffected, and the August measurement + records that name `backend-tornado` are kept as written -- they are what was measured then. + ### Added - `backend-cuda` is now published. It ships one `sm_80` PTX module serving every device of compute @@ -45,8 +66,11 @@ reading of 1.788x as a dispatch ceiling was wrong; the binding constraint was th path. **One host and one model.** G1 has no tolerance, so each qualifying hardware profile and each -architecture family earns its own run before the claim generalises. No published Java artifact -behaviour changed in this release: `v0.3.52..HEAD` touches no other published module. +architecture family earns its own run before the claim generalises. + +The published surface changes in two ways and no others: `backend-cuda` is added and +`backend-tornado` is removed. No remaining published module's Java behaviour changed -- the rest of +`v0.3.52..HEAD` touches only benchmark applications and evidence. ## [0.3.52] - 2026-10-06 diff --git a/README.md b/README.md index a925e188..e97f9324 100644 --- a/README.md +++ b/README.md @@ -34,9 +34,10 @@ tokens. Models implements that pipeline on Java 25 and uses the Vector API for CPU SIMD execution: - `backend-java` executes every inference kernel in Java. -- `backend-tornado` optionally compiles the Java Q4_0 projection kernels for a - qualified NVIDIA GPU. It keeps the Models graph in-process and falls back to - the Vector API when the device or artifact is not eligible. +- `backend-cuda` optionally runs the K-quant projections and grouped-query decode + attention on a qualified NVIDIA GPU, through our own Rust kernels compiled to + PTX. It keeps the Models graph in-process and falls back to the Vector API when + the device or artifact is not eligible. - `backend-native` runs the same Java 25 and Vector API pipeline, substituting only selected, measured bottleneck kernels with a small Models-owned Rust library through Java's Foreign Function and Memory (FFM) API. @@ -212,9 +213,11 @@ dependencies { } ``` -For qualified NVIDIA acceleration, add `backend-tornado` and launch with a -matching TornadoVM PTX runtime. The default loader performs eager readiness and -uses the Vector API when the GPU cannot safely retain the compiled plans. See +For qualified NVIDIA acceleration, add `backend-cuda`. One `sm_80` PTX module +ships inside the jar for every device of compute capability 8.0 or above, so no +external runtime is installed; the path is opt-in through +`CudaGgufBatchedMatrixKernel.open()` and falls back to the Vector API when the +device or artifact is not eligible. See [Java GPU acceleration](https://integrallis.github.io/models/docs/models/current/gpu-acceleration.html). Use Apple's on-device system model on a supported Apple Silicon Mac: @@ -375,7 +378,7 @@ documented in [Execution planning](https://integrallis.github.io/models/docs/mod | Model routing | `models-router` | adaptive selection and failover across in-process and hosted clients, with hard per-request capability and data-boundary requirements, without a provider SDK dependency | | Vector storage | `models-embedding` | optional bridge to `vectors` | | Apple on-device model | `backend-apple` | Apple Foundation Models through Java FFM | -| Java GPU acceleration | `backend-tornado` | optional Java-authored Q4_0 projections on qualified NVIDIA GPUs | +| NVIDIA GPU acceleration | `backend-cuda` | optional Rust-authored PTX K-quant projections and decode attention | These adapters are implemented and tested against the same backend contracts; they do not select hidden inference paths. Their framework dependencies are @@ -409,7 +412,7 @@ RAG, Javadocs, and release testing. - [Executable Java notebooks](notebooks/README.md) - [Apple Foundation Models bridge](models-backend-apple/README.md) -- [Java GPU acceleration](backend-tornado/README.md) +- [NVIDIA GPU acceleration](backend-cuda/README.md) - [Native kernel backend](backend-native/README.md) ## Build diff --git a/RELEASING.md b/RELEASING.md index b92a0b50..990877da 100644 --- a/RELEASING.md +++ b/RELEASING.md @@ -5,7 +5,7 @@ JReleaser signs and validates one Maven Central bundle, and the workflow creates the GitHub release. The publication allowlist contains `models-api`, `models-runtime`, `models`, -`models-rag`, `models-semantic-order`, `backend-java`, `backend-tornado`, `backend-native`, +`models-rag`, `models-semantic-order`, `backend-java`, `backend-native`, `backend-cuda`, `backend-apple`, `models-langchain4j`, `models-spring-ai`, `models-spring-boot-starter`, `models-embedding`, `models-audio`, `models-router`, and `models-decisions`. Benchmark applications, documentation tooling, and modules containing only package scaffolding are not @@ -36,10 +36,12 @@ can opt into it and read the gates. API numeric kernels. The release workflow builds and tests the Models-owned Rust kernels on every supported native platform and compiles the Apple Foundation Models bridge on macOS before staging the signed Maven artifacts. -`backend-tornado` is an optional JVM artifact. The hosted release workflow verifies its Java -fallback and publication shape. Before release, run the public loader and exact CPU/GPU output -parity gate on each qualified NVIDIA hardware profile and retain the measurements under -`models-accelerator-bench/results/`. +Models ships **one** GPU implementation, `backend-cuda`. The TornadoVM arm was removed in 0.3.53: +measured on the same RTX 4090 and the same model, `backend-cuda` reached exact token parity and +31.26 tok/s decode where TornadoVM managed 6.49 tok/s over 36 failed `cuLaunchKernel` calls, and it +required a separately installed runtime that the Maven artifact could not carry. The release +precondition that pointed at `models-accelerator-bench/results/` went with it; `backend-cuda`'s +gates are `cuda-kernel-gate --mode capability|parity|decode` and need no external runtime. The workflow uses the same Maven Central and GPG secrets as `mfcqi-java`: `MAVENCENTRAL_USERNAME`, `MAVENCENTRAL_PASSWORD`, `GPG_PUBLIC_KEY`, diff --git a/backend-tornado/.github/badges/mfcqi.json b/backend-tornado/.github/badges/mfcqi.json deleted file mode 100644 index 4b6b9094..00000000 --- a/backend-tornado/.github/badges/mfcqi.json +++ /dev/null @@ -1,7 +0,0 @@ -{ - "schemaVersion": 1, - "label": "MFCQI", - "message": "0.83 (excellent)", - "color": "brightgreen", - "cacheSeconds": 3600 -} \ No newline at end of file diff --git a/backend-tornado/README.md b/backend-tornado/README.md deleted file mode 100644 index 18085844..00000000 --- a/backend-tornado/README.md +++ /dev/null @@ -1,35 +0,0 @@ -# Java GPU acceleration - -`backend-tornado` is an optional in-process device backend for Models. Its kernels are Java source; -TornadoVM compiles eligible Q4_0, Q4_K and Q6_K projections for a GPU at runtime. It does not invoke -an external model server or embed another inference engine. - -The Maven artifact does not bundle or transitively install TornadoVM's device runtime. Applications -provide a compatible TornadoVM distribution and launch configuration explicitly; without it, the -default loader uses the Java Vector API. - -Adding this module also registers its accelerator through Java `ServiceLoader`. -`PureJavaBackend.loadAutomatic(...)`—and ModelJars' Java backend path—select it only when the exact -artifact and device pass the capacity gate. - -The qualified production scope is deliberately narrow: - -- NVIDIA GPUs reached through TornadoVM's PTX backend; -- Q4_0 (Q8_0 activations) and Q4_K / Q6_K (Q8_K activations) GGUF projection work for both prefill - and single-token decode, with the two activation families never mixed inside one grouped dispatch; -- a single weight tensor under 2 GiB, the limit of TornadoVM's 32-bit array index; -- attention and unsupported tensor formats remain on the Java Vector API; and -- eager readiness compiles reusable plans before the first visible request. - -The capacity selector passed exact output-parity and full-model gates on NVIDIA A16 and A40 profiles -with Q4_0. The K-quant kernels are covered off-device against the vectors-core CPU kernels — exactly -on exactly-representable super-blocks, and within two float roundings per super-block otherwise — and -have not yet been run on a GPU. AMD, Intel, and Metal devices remain on the CPU fallback until they -pass equivalent real hardware gates. - -Run the hardware gate with the `accelerator-profile` command of `models-bench`, launched through the -TornadoVM launcher (see the models-bench README). It reports the selected device, readiness time, -prefill and decode throughput, and how many projections reached the device per GGUF weight format. - -See the published [Java GPU acceleration guide](https://integrallis.github.io/models/docs/models/current/gpu-acceleration.html) -for dependencies, launcher requirements, status reporting, and measured evidence. diff --git a/backend-tornado/build.gradle.kts b/backend-tornado/build.gradle.kts deleted file mode 100644 index aebb5933..00000000 --- a/backend-tornado/build.gradle.kts +++ /dev/null @@ -1,55 +0,0 @@ -import org.gradle.testing.jacoco.tasks.JacocoCoverageVerification -import org.gradle.testing.jacoco.tasks.JacocoReport - -description = "Optional Java-authored GPU acceleration for the Models pure-Java backend" - -dependencies { - api(project(":backend-java")) - compileOnly("io.github.beehive-lab:tornado-api:5.2.0-jdk25") - testImplementation("io.github.beehive-lab:tornado-api:5.2.0-jdk25") - testRuntimeOnly("io.github.beehive-lab:tornado-runtime:5.2.0-jdk25") - testImplementation("com.integrallis:vectors-core:${providers.gradleProperty("vectorsVersion").get()}") -} - -val configuredQwenModel = providers.systemProperty("models.fixtures.qwen306BQ40") -val acceleratorRequired = providers.systemProperty("models.accelerator.required") -val acceleratorExpected = providers.systemProperty("models.accelerator.expected") - -tasks.withType().configureEach { - configuredQwenModel.orNull?.let { - systemProperty("models.fixtures.qwen306BQ40", it) - } - acceleratorRequired.orNull?.let { - systemProperty("models.accelerator.required", it) - } - acceleratorExpected.orNull?.let { - systemProperty("models.accelerator.expected", it) - } -} - -// Driver discovery, model loading, and Tornado execution plans are covered by the opt-in model -// integration test and the release hardware gates. Keep the ordinary unit-coverage denominator on -// the device-independent selection, validation, and Java kernel logic. -val hardwareIntegrationClasses = - listOf( - "**/TornadoBackend.class", - "**/TornadoBackendRuntime.class", - "**/TornadoRuntimeDevices.class", - "**/TornadoGgufBatchedMatrixKernel*.class" - ) - -tasks.withType().configureEach { - classDirectories.setFrom( - files(classDirectories.files.map { directory -> - fileTree(directory) { exclude(hardwareIntegrationClasses) } - }) - ) -} - -tasks.withType().configureEach { - classDirectories.setFrom( - files(classDirectories.files.map { directory -> - fileTree(directory) { exclude(hardwareIntegrationClasses) } - }) - ) -} diff --git a/backend-tornado/gradle.lockfile b/backend-tornado/gradle.lockfile deleted file mode 100644 index f8ef23f5..00000000 --- a/backend-tornado/gradle.lockfile +++ /dev/null @@ -1,75 +0,0 @@ -# This is a Gradle generated file for dependency locking. -# Manual edits can break the build and are not advised. -# This file is expected to be part of source control. -biz.aQute.bnd:biz.aQute.bnd.annotation:7.1.0=compileClasspath,testCompileClasspath -com.fasterxml.jackson.core:jackson-core:2.21.7=runtimeClasspath,testRuntimeClasspath -com.fasterxml.jackson:jackson-bom:2.21.7=runtimeClasspath,testRuntimeClasspath -com.github.spotbugs:spotbugs-annotations:4.9.8=spotbugs -com.github.spotbugs:spotbugs:4.9.8=spotbugs -com.github.stephenc.jcip:jcip-annotations:1.0-1=spotbugs -com.google.code.findbugs:jsr305:3.0.2=spotbugs -com.google.code.gson:gson:2.13.2=spotbugs -com.google.errorprone:error_prone_annotations:2.38.0=compileClasspath,testCompileClasspath -com.google.errorprone:error_prone_annotations:2.41.0=spotbugs -com.integrallis:vectors-core:0.1.28=runtimeClasspath,testCompileClasspath,testRuntimeClasspath -commons-io:commons-io:2.20.0=spotbugs -io.github.beehive-lab:tornado-api:5.2.0-jdk25=compileClasspath,testCompileClasspath,testRuntimeClasspath -io.github.beehive-lab:tornado-runtime:5.2.0-jdk25=testRuntimeClasspath -jaxen:jaxen:2.0.0=spotbugs -net.bytebuddy:byte-buddy-agent:1.17.7=testCompileClasspath,testRuntimeClasspath -net.bytebuddy:byte-buddy:1.17.7=testCompileClasspath,testRuntimeClasspath -net.sf.jopt-simple:jopt-simple:4.6=compileClasspath,testCompileClasspath,testRuntimeClasspath -net.sf.saxon:Saxon-HE:12.9=spotbugs -org.apache.bcel:bcel:6.11.0=spotbugs -org.apache.commons:commons-lang3:3.19.0=spotbugs -org.apache.commons:commons-math3:3.2=compileClasspath,testCompileClasspath,testRuntimeClasspath -org.apache.commons:commons-text:1.14.0=spotbugs -org.apache.logging.log4j:log4j-api:2.25.2=spotbugs -org.apache.logging.log4j:log4j-api:2.25.4=compileClasspath,testCompileClasspath,testRuntimeClasspath -org.apache.logging.log4j:log4j-core:2.25.2=spotbugs -org.apache.logging.log4j:log4j-core:2.25.4=compileClasspath,testCompileClasspath,testRuntimeClasspath -org.apiguardian:apiguardian-api:1.1.2=testCompileClasspath -org.assertj:assertj-core:3.27.2=testCompileClasspath,testRuntimeClasspath -org.dom4j:dom4j:2.2.0=spotbugs -org.graalvm.polyglot:polyglot:25.0.2=compileClasspath,testCompileClasspath,testRuntimeClasspath -org.graalvm.sdk:collections:25.0.2=compileClasspath,testCompileClasspath,testRuntimeClasspath -org.graalvm.sdk:nativeimage:25.0.2=compileClasspath,testCompileClasspath,testRuntimeClasspath -org.graalvm.sdk:word:25.0.2=compileClasspath,testCompileClasspath,testRuntimeClasspath -org.jacoco:org.jacoco.agent:0.8.14=jacocoAgent,jacocoAnt -org.jacoco:org.jacoco.ant:0.8.14=jacocoAnt -org.jacoco:org.jacoco.core:0.8.14=jacocoAnt -org.jacoco:org.jacoco.report:0.8.14=jacocoAnt -org.jspecify:jspecify:1.0.0=compileClasspath,testCompileClasspath -org.junit.jupiter:junit-jupiter-api:5.11.4=testCompileClasspath -org.junit.jupiter:junit-jupiter-api:5.13.4=testRuntimeClasspath -org.junit.jupiter:junit-jupiter-engine:5.13.4=testRuntimeClasspath -org.junit.jupiter:junit-jupiter-params:5.11.4=testCompileClasspath -org.junit.jupiter:junit-jupiter-params:5.13.4=testRuntimeClasspath -org.junit.jupiter:junit-jupiter:5.11.4=testCompileClasspath -org.junit.jupiter:junit-jupiter:5.13.4=testRuntimeClasspath -org.junit.platform:junit-platform-commons:1.11.4=testCompileClasspath -org.junit.platform:junit-platform-commons:1.13.4=testRuntimeClasspath -org.junit.platform:junit-platform-engine:1.13.4=testRuntimeClasspath -org.junit.platform:junit-platform-launcher:1.13.4=testRuntimeClasspath -org.junit:junit-bom:5.11.4=testCompileClasspath -org.junit:junit-bom:5.13.4=testRuntimeClasspath -org.junit:junit-bom:5.14.0=spotbugs -org.mockito:mockito-core:5.23.0=testCompileClasspath,testRuntimeClasspath -org.mockito:mockito-junit-jupiter:5.23.0=testCompileClasspath,testRuntimeClasspath -org.objenesis:objenesis:3.3=testRuntimeClasspath -org.openjdk.jmh:jmh-core:1.29=compileClasspath,testCompileClasspath,testRuntimeClasspath -org.opentest4j:opentest4j:1.3.0=testCompileClasspath,testRuntimeClasspath -org.osgi:org.osgi.annotation.bundle:2.0.0=compileClasspath,testCompileClasspath -org.osgi:org.osgi.annotation.versioning:1.1.2=compileClasspath,testCompileClasspath -org.osgi:org.osgi.resource:1.0.0=compileClasspath,testCompileClasspath -org.osgi:org.osgi.service.serviceloader:1.0.0=compileClasspath,testCompileClasspath -org.ow2.asm:asm-analysis:9.9=spotbugs -org.ow2.asm:asm-commons:9.9=jacocoAnt,spotbugs -org.ow2.asm:asm-tree:9.9=jacocoAnt,spotbugs -org.ow2.asm:asm-util:9.9=spotbugs -org.ow2.asm:asm:9.9=jacocoAnt,spotbugs -org.slf4j:slf4j-api:2.0.17=runtimeClasspath,spotbugs,spotbugsSlf4j,testRuntimeClasspath -org.slf4j:slf4j-simple:2.0.17=spotbugsSlf4j -org.snmp4j:snmp4j:2.8.6=testRuntimeClasspath -org.xmlresolver:xmlresolver:5.3.3=spotbugs -empty=annotationProcessor,cyclonedxBom,spotbugsPlugins,testAnnotationProcessor diff --git a/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/AcceleratorEligibility.java b/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/AcceleratorEligibility.java deleted file mode 100644 index da178887..00000000 --- a/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/AcceleratorEligibility.java +++ /dev/null @@ -1,320 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.backend.tornado; - -import java.time.Duration; -import java.util.Comparator; -import java.util.List; -import java.util.Locale; -import java.util.Objects; - -/** - * Device-vendor and capacity policy applied before constructing accelerator execution plans. - * - *

The gate adds up what a model would actually place on the device — retained weights under the - * plan shape the kernel builds, per-plan scratch, and any device-resident KV cache — and compares - * it against the device's global memory less a safety margin. Where a term is unknown it is refused - * rather than assumed to be zero: a model whose weights exceed {@link - * #COARSE_PROFILE_WEIGHT_LIMIT_BYTES} is not admitted on a file-size-only budget, because at that - * size the unknown terms are larger than the margin. - */ -public final class AcceleratorEligibility { - - /** - * Flat allowance for the TornadoVM device context, compiled code, and driver bookkeeping. - * - *

Unmeasured. It is a constant carried over from the published selector, kept so the admitted - * budget for the qualified A16-2Q and A40-4Q profiles is unchanged. Calibrating it needs a GPU - * host: {@code TornadoExecutionPlan.getCurrentDeviceMemoryUsage()} reports real usage after - * readiness, and until that is run this term must not be described as measured. - */ - static final long BASE_PLAN_OVERHEAD_BYTES = 256L * 1024L * 1024L; - - /** Maximum single allocation assumed for a coarse, file-size-only request. */ - static final long COARSE_MAX_SINGLE_ALLOCATION_BYTES = 512L * 1024L * 1024L; - - /** - * Hard cap on one uploaded tensor. - * - *

{@code ByteArray} stores its element count in an {@code int} and {@code - * ByteArray.fromSegment} computes it as {@code (int) segment.byteSize()}, so a tensor at or above - * 2 GiB is truncated on construction and the following {@code MemorySegment.copy} of the full - * tensor fails. Read from {@code tornado-api 5.2.0-jdk25}. - */ - static final long MAX_TORNADO_ARRAY_BYTES = Integer.MAX_VALUE; - - /** - * Weight ceiling for admitting a model on a coarse, file-size-only request. - * - *

Below this a model's KV cache, per-plan scratch, and largest tensor all fit comfortably - * inside {@link #BASE_PLAN_OVERHEAD_BYTES}; above it they do not, and admitting the model would - * mean gating on a number known to be incomplete. - */ - static final long COARSE_PROFILE_WEIGHT_LIMIT_BYTES = 8L * 1024L * 1024L * 1024L; - - /** - * Eager-readiness ceiling above which a configuration is not deployable. - * - *

Fixed in advance by gate G5 of the large-model GPU campaign pre-registration. - */ - static final Duration READINESS_BUDGET = Duration.ofSeconds(120); - - /** Fraction of device global memory the gate refuses to plan into. */ - private static final long SAFETY_DIVISOR = 4L; - - private AcceleratorEligibility() {} - - /** - * Selects a device using the coarse, file-size-only budget available before a model is parsed. - * - * @param devices devices reported by the TornadoVM runtime - * @param modelSizeBytes size of the GGUF file on disk - * @param accelerateDecode whether separate single-token decode plans are retained - */ - public static Decision select( - List devices, long modelSizeBytes, boolean accelerateDecode) { - Objects.requireNonNull(devices, "devices"); - if (modelSizeBytes <= 0) { - return Decision.ineligible("model size must be positive", 0, null); - } - return select( - devices, DeviceMemoryRequest.ofModelFile("model", modelSizeBytes, accelerateDecode)); - } - - /** Selects a device for a fully specified device-memory request. */ - public static Decision select(List devices, DeviceMemoryRequest request) { - Objects.requireNonNull(devices, "devices"); - Objects.requireNonNull(request, "request"); - - DeviceBudget shipped = budget(request, PlanShapeStrategy.PER_SHAPE_WHOLE_MODEL); - - List qualified = - devices.stream() - .filter(device -> "PTX".equals(device.backend()) || "CUDA".equals(device.backend())) - .filter(device -> "GPU".equals(device.type())) - .sorted(Comparator.comparingLong(DeviceCapabilities::globalMemoryBytes).reversed()) - .toList(); - if (qualified.isEmpty()) { - return Decision.ineligible( - "no qualified NVIDIA GPU backend was discovered; needed a PTX or CUDA backend on a GPU" - + " device, found " - + describe(devices), - shipped.totalBytes(), - shipped); - } - - if (request.largestAllocationBytes() >= MAX_TORNADO_ARRAY_BYTES) { - return Decision.ineligible( - "largest device allocation is " - + DeviceBudget.gib(request.largestAllocationBytes()) - + " but a TornadoVM ByteArray holds at most " - + DeviceBudget.gib(MAX_TORNADO_ARRAY_BYTES) - + " (int element count); the tensor must be split before it can be uploaded", - shipped.totalBytes(), - shipped); - } - - if (!request.detailed() && request.weightBytes() > COARSE_PROFILE_WEIGHT_LIMIT_BYTES) { - return Decision.ineligible( - "model '" - + request.modelLabel() - + "' has " - + DeviceBudget.gib(request.weightBytes()) - + " of weights, above the " - + DeviceBudget.gib(COARSE_PROFILE_WEIGHT_LIMIT_BYTES) - + " limit for a file-size-only device budget; KV residency, retained-plan count," - + " per-plan scratch and largest-tensor size are all unknown in this request." - + " Build a DeviceMemoryRequest.detailed(...) from the model's tensor inventory" - + " before admitting a model this size", - shipped.totalBytes(), - shipped); - } - - DeviceCapabilities best = qualified.get(0); - for (DeviceCapabilities device : qualified) { - long safeCapacity = safeCapacity(device); - if (shipped.totalBytes() <= safeCapacity && allocationFits(request, device)) { - Duration readiness = shipped.estimatedPlanCompileTime(); - if (readiness.compareTo(READINESS_BUDGET) > 0) { - return Decision.ineligible( - "device " - + device.name() - + " fits the plan but the " - + request.retainedPlanCount() - + " retained plans would need about " - + seconds(readiness) - + " of eager readiness at the measured " - + String.format(Locale.ROOT, "%.1f", DeviceBudget.PLANS_COMPILED_PER_SECOND) - + " plans/s (A40-4Q, 223 plans in 13.554 s), above the " - + seconds(READINESS_BUDGET) - + " deployable ceiling; reduce the retained plan count before admitting this" - + " model", - shipped.totalBytes(), - shipped); - } - return new Decision( - true, - device, - "eligible", - shipped.totalBytes(), - shipped, - PlanShapeStrategy.PER_SHAPE_WHOLE_MODEL); - } - } - - for (DeviceCapabilities device : qualified) { - if (!allocationFits(request, device)) { - continue; - } - long safeCapacity = safeCapacity(device); - for (PlanShapeStrategy strategy : PlanShapeStrategy.values()) { - if (strategy.implemented() || !strategy.capacityComputable()) { - continue; - } - DeviceBudget alternative = budget(request, strategy); - if (alternative.totalBytes() <= safeCapacity) { - return Decision.ineligible( - "device " - + device.name() - + " has " - + DeviceBudget.gib(safeCapacity) - + " usable of " - + DeviceBudget.gib(device.globalMemoryBytes()) - + " but the shipped plan shape needs " - + DeviceBudget.gib(shipped.totalBytes()) - + " (" - + shipped.describe() - + "); the model would fit in " - + DeviceBudget.gib(alternative.totalBytes()) - + " under " - + strategy - + ", which is not implemented: " - + strategy.limitation(), - shipped.totalBytes(), - shipped); - } - } - } - - DeviceCapabilities target = best; - long safeCapacity = safeCapacity(target); - if (!allocationFits(request, target)) { - return Decision.ineligible( - "device " - + target.name() - + " allows a single allocation of at most " - + DeviceBudget.gib(target.maxAllocationBytes()) - + " but the plan needs one of " - + DeviceBudget.gib(largestAllocation(request)), - shipped.totalBytes(), - shipped); - } - return Decision.ineligible( - "device " - + target.name() - + " has " - + DeviceBudget.gib(safeCapacity) - + " usable of " - + DeviceBudget.gib(target.globalMemoryBytes()) - + " but the plan needs " - + DeviceBudget.gib(shipped.totalBytes()) - + " (" - + shipped.describe() - + "), short by " - + DeviceBudget.gib(shipped.totalBytes() - safeCapacity) - + "; the same plans also hold " - + DeviceBudget.gib(shipped.hostWeightCopyBytes()) - + " of host weight copies", - shipped.totalBytes(), - shipped); - } - - /** Computes the device budget a request costs under one plan shape. */ - public static DeviceBudget budget(DeviceMemoryRequest request, PlanShapeStrategy strategy) { - return DeviceBudget.of(request, strategy, BASE_PLAN_OVERHEAD_BYTES); - } - - private static long safeCapacity(DeviceCapabilities device) { - return device.globalMemoryBytes() - device.globalMemoryBytes() / SAFETY_DIVISOR; - } - - private static long largestAllocation(DeviceMemoryRequest request) { - return request.detailed() - ? request.largestAllocationBytes() - : Math.min(request.weightBytes(), COARSE_MAX_SINGLE_ALLOCATION_BYTES); - } - - private static boolean allocationFits(DeviceMemoryRequest request, DeviceCapabilities device) { - return largestAllocation(request) <= device.maxAllocationBytes(); - } - - private static String describe(List devices) { - if (devices.isEmpty()) { - return "no devices"; - } - StringBuilder text = new StringBuilder(64); - for (DeviceCapabilities device : devices) { - if (text.length() > 0) { - text.append(", "); - } - text.append(device.name()) - .append(" [") - .append(device.backend()) - .append('/') - .append(device.type()) - .append(']'); - } - return text.toString(); - } - - private static String seconds(Duration duration) { - return String.format(Locale.ROOT, "%.0f s", duration.toMillis() / 1000.0); - } - - /** One device as reported by the TornadoVM runtime. */ - public record DeviceCapabilities( - String name, String backend, String type, long globalMemoryBytes, long maxAllocationBytes) { - public DeviceCapabilities { - Objects.requireNonNull(name, "name"); - Objects.requireNonNull(backend, "backend"); - Objects.requireNonNull(type, "type"); - } - } - - /** - * The outcome of device selection. - * - * @param eligible whether a device was admitted - * @param device the admitted device, or {@code null} - * @param reason {@code "eligible"}, or a message naming what was needed and what was available - * @param requiredBytes device bytes the shipped plan shape would need - * @param budget the itemised budget behind {@code requiredBytes}, or {@code null} when the - * request was rejected before one could be formed - * @param strategy the plan shape the admitted device was admitted under, or {@code null} - */ - public record Decision( - boolean eligible, - DeviceCapabilities device, - String reason, - long requiredBytes, - DeviceBudget budget, - PlanShapeStrategy strategy) { - - private static Decision ineligible(String reason, long requiredBytes, DeviceBudget budget) { - return new Decision(false, null, reason, requiredBytes, budget, null); - } - } -} diff --git a/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/DeviceBudget.java b/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/DeviceBudget.java deleted file mode 100644 index ff87a937..00000000 --- a/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/DeviceBudget.java +++ /dev/null @@ -1,101 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.backend.tornado; - -import java.time.Duration; -import java.util.Locale; -import java.util.Objects; - -/** - * The device and host bytes one {@link DeviceMemoryRequest} costs under one {@link - * PlanShapeStrategy}, itemised so an ineligibility message can name every term. - * - * @param strategy the plan shape this budget was computed for - * @param deviceWeightBytes model weights retained on the device - * @param planScratchBytes device activation, scale, and output buffers across all retained plans - * @param kvCacheBytes KV cache bytes retained on the device - * @param baseOverheadBytes flat runtime allowance for the TornadoVM context and compiled code - * @param totalBytes sum of the device terms - * @param hostWeightCopyBytes off-heap host copies the plans hold, on top of the mapped GGUF - * @param estimatedPlanCompileTime eager-readiness estimate from the measured plan compile rate - */ -public record DeviceBudget( - PlanShapeStrategy strategy, - long deviceWeightBytes, - long planScratchBytes, - long kvCacheBytes, - long baseOverheadBytes, - long totalBytes, - long hostWeightCopyBytes, - Duration estimatedPlanCompileTime) { - - /** - * Plans compiled per second during eager readiness. - * - *

Measured, not assumed: the Vultr A40-4Q gate on 2026-08-29 compiled 223 retained plans in - * 13.554 s ({@code models-accelerator-bench/results/vultr-a40-4q-2026-08-29.md}). The A16-2Q gate - * on the same artifact took 14.278 s for the same plan set, so this rate is the faster of the two - * measured profiles and the readiness estimate it produces is optimistic. - */ - static final double PLANS_COMPILED_PER_SECOND = 223.0 / 13.554; - - public DeviceBudget { - Objects.requireNonNull(strategy, "strategy"); - Objects.requireNonNull(estimatedPlanCompileTime, "estimatedPlanCompileTime"); - } - - static DeviceBudget of(DeviceMemoryRequest request, PlanShapeStrategy strategy, long baseBytes) { - Objects.requireNonNull(request, "request"); - Objects.requireNonNull(strategy, "strategy"); - long weights = strategy.deviceWeightBytes(request); - long scratch = request.planScratchBytes(); - long kv = request.deviceKvCacheBytes(); - long total = Math.addExact(Math.addExact(weights, scratch), Math.addExact(kv, baseBytes)); - Duration readiness = - Duration.ofMillis( - Math.round(request.retainedPlanCount() / PLANS_COMPILED_PER_SECOND * 1000.0)); - return new DeviceBudget( - strategy, - weights, - scratch, - kv, - baseBytes, - total, - strategy.hostWeightCopyBytes(request), - readiness); - } - - /** A one-line itemisation suitable for an operator-facing message. */ - public String describe() { - StringBuilder text = new StringBuilder(128); - text.append("weights ") - .append(gib(deviceWeightBytes)) - .append(" (") - .append(strategy) - .append("), KV ") - .append(gib(kvCacheBytes)) - .append(", plan scratch ") - .append(gib(planScratchBytes)) - .append(", base ") - .append(gib(baseOverheadBytes)); - return text.toString(); - } - - /** Formats a byte count as GiB with two decimals. */ - public static String gib(long bytes) { - return String.format(Locale.ROOT, "%.2f GiB", bytes / (double) (1024L * 1024L * 1024L)); - } -} diff --git a/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/DeviceMemoryRequest.java b/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/DeviceMemoryRequest.java deleted file mode 100644 index 8cb69c78..00000000 --- a/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/DeviceMemoryRequest.java +++ /dev/null @@ -1,179 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.backend.tornado; - -import java.util.Objects; - -/** - * What one model would place on an accelerator, in the terms the capacity gate can add up. - * - *

A request is either coarse or detailed. A coarse request is derived from the - * GGUF file size alone: it knows the upper bound on uploaded weights and nothing else, so retained - * plan count, largest single tensor, and device KV residency are all reported as unknown rather - * than silently assumed to be zero. A detailed request is built from the tensor inventory and - * carries all of those terms. - * - * @param modelLabel human-readable model identity used in ineligibility messages - * @param weightBytes bytes of model weights that would be uploaded once, per retained shape - * @param retainedShapes distinct batch shapes whose plans are held at once (prefill, or prefill and - * decode) - * @param retainedPlanCount distinct {@code (weight, shape)} plans that would be retained; {@code 0} - * means unknown - * @param planScratchBytes device activation, scale, and output buffers summed over every retained - * plan; {@code 0} means unknown - * @param largestAllocationBytes largest single device allocation the plans would ask for, whether a - * weight tensor or a scratch buffer; {@code 0} means unknown - * @param deviceKvCacheBytes KV cache bytes that would live on the device at the planned context - * length; {@code 0} means the KV cache stays in host memory - * @param detailed whether the weight, plan, and tensor terms came from a real tensor inventory - */ -public record DeviceMemoryRequest( - String modelLabel, - long weightBytes, - int retainedShapes, - int retainedPlanCount, - long planScratchBytes, - long largestAllocationBytes, - long deviceKvCacheBytes, - boolean detailed) { - - public DeviceMemoryRequest { - modelLabel = requireLabel(modelLabel); - requirePositive("weightBytes", weightBytes); - if (retainedShapes < 1) { - throw new IllegalArgumentException("retainedShapes must be at least 1: " + retainedShapes); - } - requireNotNegative("retainedPlanCount", retainedPlanCount); - requireNotNegative("planScratchBytes", planScratchBytes); - requireNotNegative("largestAllocationBytes", largestAllocationBytes); - requireNotNegative("deviceKvCacheBytes", deviceKvCacheBytes); - } - - /** - * Builds the coarse request the loader can form before parsing a model: the whole GGUF file is - * treated as uploadable weights and nothing else is known. - * - *

This is what {@code TornadoBackend.open} has available, and it reproduces the budget the - * published A16-2Q and A40-4Q qualification runs were admitted under. - */ - public static DeviceMemoryRequest ofModelFile( - String modelLabel, long fileBytes, boolean accelerateDecode) { - return new DeviceMemoryRequest( - modelLabel, fileBytes, accelerateDecode ? 2 : 1, 0, 0L, 0L, 0L, false); - } - - /** Starts a detailed request built from a real tensor inventory. */ - public static Builder detailed(String modelLabel) { - return new Builder(modelLabel); - } - - /** Returns this request with a device-resident KV cache of the given size. */ - public DeviceMemoryRequest withDeviceKvCacheBytes(long bytes) { - return new DeviceMemoryRequest( - modelLabel, - weightBytes, - retainedShapes, - retainedPlanCount, - planScratchBytes, - largestAllocationBytes, - bytes, - detailed); - } - - /** Mutable assembly for a detailed request. */ - public static final class Builder { - private final String modelLabel; - private long weightBytes; - private int retainedShapes = 1; - private int retainedPlanCount; - private long planScratchBytes; - private long largestAllocationBytes; - private long deviceKvCacheBytes; - - private Builder(String modelLabel) { - this.modelLabel = requireLabel(modelLabel); - } - - /** Sets the bytes of weights uploaded once, per retained shape. */ - public Builder weightBytes(long bytes) { - this.weightBytes = bytes; - return this; - } - - /** Sets how many batch shapes are retained at once. */ - public Builder retainedShapes(int shapes) { - this.retainedShapes = shapes; - return this; - } - - /** Sets how many distinct {@code (weight, shape)} plans are retained. */ - public Builder retainedPlanCount(int plans) { - this.retainedPlanCount = plans; - return this; - } - - /** Sets the device activation, scale, and output buffers summed over every retained plan. */ - public Builder planScratchBytes(long bytes) { - this.planScratchBytes = bytes; - return this; - } - - /** Sets the largest single device allocation the plans would ask for. */ - public Builder largestAllocationBytes(long bytes) { - this.largestAllocationBytes = bytes; - return this; - } - - /** Sets the KV cache bytes that would live on the device at the planned context length. */ - public Builder deviceKvCacheBytes(long bytes) { - this.deviceKvCacheBytes = bytes; - return this; - } - - /** Builds the detailed request. */ - public DeviceMemoryRequest build() { - return new DeviceMemoryRequest( - modelLabel, - weightBytes, - retainedShapes, - retainedPlanCount, - planScratchBytes, - largestAllocationBytes, - deviceKvCacheBytes, - true); - } - } - - private static String requireLabel(String value) { - Objects.requireNonNull(value, "modelLabel"); - if (value.isBlank()) { - throw new IllegalArgumentException("modelLabel must not be blank"); - } - return value.strip(); - } - - private static void requirePositive(String name, long value) { - if (value <= 0) { - throw new IllegalArgumentException(name + " must be positive: " + value); - } - } - - private static void requireNotNegative(String name, long value) { - if (value < 0) { - throw new IllegalArgumentException(name + " must not be negative: " + value); - } - } -} diff --git a/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/KQuantProjectionKernel.java b/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/KQuantProjectionKernel.java deleted file mode 100644 index 29cf10ab..00000000 --- a/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/KQuantProjectionKernel.java +++ /dev/null @@ -1,568 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.backend.tornado; - -import uk.ac.manchester.tornado.api.annotations.Parallel; -import uk.ac.manchester.tornado.api.types.arrays.ByteArray; -import uk.ac.manchester.tornado.api.types.arrays.FloatArray; -import uk.ac.manchester.tornado.api.types.arrays.IntArray; - -/** - * Java-authored TornadoVM kernels for production-compatible K-quant by Q8_K projections. - * - *

Q4_K_M catalog models mix two super-block formats in one tensor set: Q4_K for most projections - * and Q6_K for the value and a share of the down projections. Both consume the same Q8_K activation - * preparation, so one host-side quantization feeds every kernel here. - * - *

GGUF super-block layouts implemented here (little-endian, 256 values per super-block): - * - *

    - *
  • Q4_K, 144 bytes: {@code d} (fp16) at 0, {@code dmin} (fp16) at 2, twelve bytes of - * six-bit packed scales and mins at 4, and 128 bytes of 4-bit quants at 16. The 256 values - * are eight groups of 32; group {@code g} reads the 32 bytes at {@code 16 + (g >>> 1) * 32} - * and takes the low nibble when {@code g} is even and the high nibble when it is odd. Group - * values are unsigned 0..15 and the offset is carried by the per-group six-bit minimum. - *
  • Q6_K, 210 bytes: 128 bytes of low nibbles, 64 bytes of high two-bit pairs, 16 signed - * int8 group scales, then {@code d} (fp16) at 208. Quants are {@code (ql | qh << 4) - 32}. - *
- * - *

Arithmetic contract. Every per-super-block reduction is accumulated in {@code int}, so - * the quantized dot products and the Q4_K minimum corrections are exact and independent of work - * scheduling. Only the per-super-block scale application is floating point, and it is applied in - * ascending super-block order with the same two-step form the vectors-core CPU kernels use ({@code - * sum + d * quantizedSum}, then {@code sum - dMin * minimumSum}). The CPU kernels fuse those two - * steps with {@code Math.fma}; this kernel uses a plain multiply and add, matching the operation - * set the existing Q4_0 kernel is known to compile to PTX with. The two agree bit-for-bit whenever - * the products are exactly representable and differ by at most one fused-multiply-add rounding per - * super-block otherwise. - */ -public final class KQuantProjectionKernel { - /** Values per K-quant super-block. */ - public static final int SUPER_BLOCK_VALUES = 256; - - /** Values per Q8_K partial-sum block; Q4_K's minimum correction reads two of these per group. */ - public static final int SUM_BLOCK_VALUES = 16; - - /** Q8_K partial-sum entries per super-block. */ - public static final int SUMS_PER_SUPER_BLOCK = SUPER_BLOCK_VALUES / SUM_BLOCK_VALUES; - - /** Bytes per Q4_K super-block. */ - public static final int Q4_K_BLOCK_BYTES = 144; - - /** Bytes per Q6_K super-block. */ - public static final int Q6_K_BLOCK_BYTES = 210; - - private static final int Q4_K_SCALES_OFFSET = 4; - private static final int Q4_K_QUANTS_OFFSET = 16; - private static final int Q4_K_GROUPS = 8; - private static final int Q4_K_GROUP_VALUES = 32; - - private static final int Q6_K_QL_BYTES = 128; - private static final int Q6_K_QH_BYTES = 64; - private static final int Q6_K_SCALE_BYTES = 16; - private static final int Q6_K_SCALE_OFFSET = Q6_K_QL_BYTES + Q6_K_QH_BYTES; - private static final int Q6_K_DELTA_OFFSET = Q6_K_SCALE_OFFSET + Q6_K_SCALE_BYTES; - - private KQuantProjectionKernel() {} - - /** Runs one work item per batch/output-row pair over host-prepared Q8_K activations. */ - public static void multiplyQ4K( - ByteArray weights, - ByteArray activations, - FloatArray activationScales, - IntArray activationSums, - FloatArray output, - int batchSize, - int rows, - int cols) { - int blocks = cols / SUPER_BLOCK_VALUES; - int sumsPerBatch = cols / SUM_BLOCK_VALUES; - for (@Parallel int outputIndex = 0; outputIndex < batchSize * rows; outputIndex++) { - int batch = outputIndex / rows; - int row = outputIndex - batch * rows; - output.set( - outputIndex, - q4kRowDot( - weights, - row, - activations, - batch * cols, - activationScales, - batch * blocks, - activationSums, - batch * sumsPerBatch, - blocks)); - } - } - - /** Computes two Q4_K projections in one dispatch over one prepared activation. */ - public static void multiplyQ4KDual( - ByteArray firstWeights, - int firstRows, - ByteArray secondWeights, - int secondRows, - ByteArray activations, - FloatArray activationScales, - IntArray activationSums, - FloatArray firstOutput, - FloatArray secondOutput, - int batchSize, - int cols) { - int blocks = cols / SUPER_BLOCK_VALUES; - int sumsPerBatch = cols / SUM_BLOCK_VALUES; - int combinedRows = firstRows + secondRows; - for (@Parallel int outputIndex = 0; outputIndex < batchSize * combinedRows; outputIndex++) { - int batch = outputIndex / combinedRows; - int combinedRow = outputIndex - batch * combinedRows; - if (combinedRow < firstRows) { - firstOutput.set( - batch * firstRows + combinedRow, - q4kRowDot( - firstWeights, - combinedRow, - activations, - batch * cols, - activationScales, - batch * blocks, - activationSums, - batch * sumsPerBatch, - blocks)); - } else { - int row = combinedRow - firstRows; - secondOutput.set( - batch * secondRows + row, - q4kRowDot( - secondWeights, - row, - activations, - batch * cols, - activationScales, - batch * blocks, - activationSums, - batch * sumsPerBatch, - blocks)); - } - } - } - - /** Computes three Q4_K projections in one dispatch over one prepared activation. */ - public static void multiplyQ4KTriple( - ByteArray firstWeights, - int firstRows, - ByteArray secondWeights, - int secondRows, - ByteArray thirdWeights, - int thirdRows, - ByteArray activations, - FloatArray activationScales, - IntArray activationSums, - FloatArray firstOutput, - FloatArray secondOutput, - FloatArray thirdOutput, - int batchSize, - int cols) { - int blocks = cols / SUPER_BLOCK_VALUES; - int sumsPerBatch = cols / SUM_BLOCK_VALUES; - int firstAndSecondRows = firstRows + secondRows; - int combinedRows = firstAndSecondRows + thirdRows; - for (@Parallel int outputIndex = 0; outputIndex < batchSize * combinedRows; outputIndex++) { - int batch = outputIndex / combinedRows; - int combinedRow = outputIndex - batch * combinedRows; - if (combinedRow < firstRows) { - firstOutput.set( - batch * firstRows + combinedRow, - q4kRowDot( - firstWeights, - combinedRow, - activations, - batch * cols, - activationScales, - batch * blocks, - activationSums, - batch * sumsPerBatch, - blocks)); - } else if (combinedRow < firstAndSecondRows) { - int row = combinedRow - firstRows; - secondOutput.set( - batch * secondRows + row, - q4kRowDot( - secondWeights, - row, - activations, - batch * cols, - activationScales, - batch * blocks, - activationSums, - batch * sumsPerBatch, - blocks)); - } else { - int row = combinedRow - firstAndSecondRows; - thirdOutput.set( - batch * thirdRows + row, - q4kRowDot( - thirdWeights, - row, - activations, - batch * cols, - activationScales, - batch * blocks, - activationSums, - batch * sumsPerBatch, - blocks)); - } - } - } - - /** Runs one work item per batch/output-row pair over host-prepared Q8_K activations. */ - public static void multiplyQ6K( - ByteArray weights, - ByteArray activations, - FloatArray activationScales, - FloatArray output, - int batchSize, - int rows, - int cols) { - int blocks = cols / SUPER_BLOCK_VALUES; - for (@Parallel int outputIndex = 0; outputIndex < batchSize * rows; outputIndex++) { - int batch = outputIndex / rows; - int row = outputIndex - batch * rows; - output.set( - outputIndex, - q6kRowDot( - weights, row, activations, batch * cols, activationScales, batch * blocks, blocks)); - } - } - - /** - * Computes the Q4_K/Q4_K/Q6_K attention projection group in one dispatch. - * - *

This is the shape a Q4_K_M model presents for query, key, and value: llama.cpp promotes the - * value projection to Q6_K while the query and key projections stay Q4_K. - */ - public static void multiplyMixedTriple( - ByteArray firstWeights, - int firstRows, - ByteArray secondWeights, - int secondRows, - ByteArray thirdWeights, - int thirdRows, - ByteArray activations, - FloatArray activationScales, - IntArray activationSums, - FloatArray firstOutput, - FloatArray secondOutput, - FloatArray thirdOutput, - int batchSize, - int cols) { - int blocks = cols / SUPER_BLOCK_VALUES; - int sumsPerBatch = cols / SUM_BLOCK_VALUES; - int firstAndSecondRows = firstRows + secondRows; - int combinedRows = firstAndSecondRows + thirdRows; - for (@Parallel int outputIndex = 0; outputIndex < batchSize * combinedRows; outputIndex++) { - int batch = outputIndex / combinedRows; - int combinedRow = outputIndex - batch * combinedRows; - if (combinedRow < firstRows) { - firstOutput.set( - batch * firstRows + combinedRow, - q4kRowDot( - firstWeights, - combinedRow, - activations, - batch * cols, - activationScales, - batch * blocks, - activationSums, - batch * sumsPerBatch, - blocks)); - } else if (combinedRow < firstAndSecondRows) { - int row = combinedRow - firstRows; - secondOutput.set( - batch * secondRows + row, - q4kRowDot( - secondWeights, - row, - activations, - batch * cols, - activationScales, - batch * blocks, - activationSums, - batch * sumsPerBatch, - blocks)); - } else { - int row = combinedRow - firstAndSecondRows; - thirdOutput.set( - batch * thirdRows + row, - q6kRowDot( - thirdWeights, - row, - activations, - batch * cols, - activationScales, - batch * blocks, - blocks)); - } - } - } - - private static float q4kRowDot( - ByteArray weights, - int row, - ByteArray activations, - int activationOffset, - FloatArray activationScales, - int scaleOffset, - IntArray activationSums, - int sumOffset, - int blocks) { - int rowOffset = row * blocks * Q4_K_BLOCK_BYTES; - float sum = 0.0f; - for (int block = 0; block < blocks; block++) { - int blockOffset = rowOffset + block * Q4_K_BLOCK_BYTES; - float activationScale = activationScales.get(scaleOffset + block); - float delta = weights.getHalfFloat(blockOffset).getFloat32() * activationScale; - float deltaMin = weights.getHalfFloat(blockOffset + 2).getFloat32() * activationScale; - int scalesOffset = blockOffset + Q4_K_SCALES_OFFSET; - int quantsOffset = blockOffset + Q4_K_QUANTS_OFFSET; - int blockActivationOffset = activationOffset + block * SUPER_BLOCK_VALUES; - int blockSumOffset = sumOffset + block * SUMS_PER_SUPER_BLOCK; - int quantizedSum = 0; - int minimumSum = 0; - for (int group = 0; group < Q4_K_GROUPS; group++) { - int packedOffset = quantsOffset + (group >>> 1) * Q4_K_GROUP_VALUES; - int shift = (group & 1) * 4; - int groupActivationOffset = blockActivationOffset + group * Q4_K_GROUP_VALUES; - int groupDot = 0; - for (int index = 0; index < Q4_K_GROUP_VALUES; index++) { - int packed = weights.get(packedOffset + index) & 0xFF; - int quant = (packed >>> shift) & 0x0F; - groupDot += quant * activations.get(groupActivationOffset + index); - } - quantizedSum += groupScale(weights, scalesOffset, group) * groupDot; - int groupSumOffset = blockSumOffset + group * 2; - minimumSum += - groupMinimum(weights, scalesOffset, group) - * (activationSums.get(groupSumOffset) + activationSums.get(groupSumOffset + 1)); - } - sum = sum + delta * quantizedSum; - sum = sum - deltaMin * minimumSum; - } - return sum; - } - - private static float q6kRowDot( - ByteArray weights, - int row, - ByteArray activations, - int activationOffset, - FloatArray activationScales, - int scaleOffset, - int blocks) { - int rowOffset = row * blocks * Q6_K_BLOCK_BYTES; - float sum = 0.0f; - for (int block = 0; block < blocks; block++) { - int blockOffset = rowOffset + block * Q6_K_BLOCK_BYTES; - float delta = - weights.getHalfFloat(blockOffset + Q6_K_DELTA_OFFSET).getFloat32() - * activationScales.get(scaleOffset + block); - int blockActivationOffset = activationOffset + block * SUPER_BLOCK_VALUES; - int blockSum = 0; - for (int half = 0; half < 2; half++) { - int lowBase = blockOffset + half * 64; - int highBase = blockOffset + Q6_K_QL_BYTES + half * 32; - int scaleBase = blockOffset + Q6_K_SCALE_OFFSET + half * 8; - int quantBase = blockActivationOffset + half * 128; - for (int index = 0; index < 32; index++) { - int scaleIndex = index >>> 4; - int lowFirst = weights.get(lowBase + index) & 0xFF; - int lowSecond = weights.get(lowBase + 32 + index) & 0xFF; - int high = weights.get(highBase + index) & 0xFF; - int first = ((lowFirst & 0x0F) | ((high & 0x03) << 4)) - 32; - int second = ((lowSecond & 0x0F) | (((high >>> 2) & 0x03) << 4)) - 32; - int third = ((lowFirst >>> 4) | (((high >>> 4) & 0x03) << 4)) - 32; - int fourth = ((lowSecond >>> 4) | (((high >>> 6) & 0x03) << 4)) - 32; - blockSum += - weights.get(scaleBase + scaleIndex) * first * activations.get(quantBase + index) - + weights.get(scaleBase + scaleIndex + 2) - * second - * activations.get(quantBase + index + 32) - + weights.get(scaleBase + scaleIndex + 4) - * third - * activations.get(quantBase + index + 64) - + weights.get(scaleBase + scaleIndex + 6) - * fourth - * activations.get(quantBase + index + 96); - } - } - sum = sum + delta * blockSum; - } - return sum; - } - - /** Decodes one of the eight six-bit group scales packed into a Q4_K super-block header. */ - private static int groupScale(ByteArray weights, int scalesOffset, int group) { - if (group < 4) { - return weights.get(scalesOffset + group) & 0x3F; - } - int low = weights.get(scalesOffset + group + 4) & 0x0F; - int high = (weights.get(scalesOffset + group - 4) & 0xFF) >>> 6; - return low | (high << 4); - } - - /** Decodes one of the eight six-bit group minima packed into a Q4_K super-block header. */ - private static int groupMinimum(ByteArray weights, int scalesOffset, int group) { - if (group < 4) { - return weights.get(scalesOffset + group + 4) & 0x3F; - } - int low = (weights.get(scalesOffset + group + 4) & 0xFF) >>> 4; - int high = (weights.get(scalesOffset + group) & 0xFF) >>> 6; - return low | (high << 4); - } - - /** - * Reproduces the production Q8_K activation preparation. - * - *

Unlike Q8_0, the Q8_K scale is not rounded through fp16 and the block is scaled by the - * signed extremum rather than the absolute maximum. The per-16-value sums are what Q4_K's minimum - * correction consumes; they always fit in a {@code short} on the CPU side, so widening them to - * {@code int} for the device is value-preserving. - */ - public static void quantize( - float[] input, byte[] activations, float[] scales, int[] sums, int batchSize, int cols) { - int blocks = cols / SUPER_BLOCK_VALUES; - for (int batch = 0; batch < batchSize; batch++) { - int rowOffset = batch * cols; - int rowScaleOffset = batch * blocks; - int rowSumOffset = batch * (cols / SUM_BLOCK_VALUES); - for (int block = 0; block < blocks; block++) { - int offset = rowOffset + block * SUPER_BLOCK_VALUES; - float extremum = 0.0f; - float absoluteMax = 0.0f; - for (int index = 0; index < SUPER_BLOCK_VALUES; index++) { - float value = input[offset + index]; - float absolute = Math.abs(value); - if (absolute > absoluteMax) { - absoluteMax = absolute; - extremum = value; - } - } - int blockSumOffset = rowSumOffset + block * SUMS_PER_SUPER_BLOCK; - if (absoluteMax == 0.0f) { - for (int index = 0; index < SUPER_BLOCK_VALUES; index++) { - activations[offset + index] = 0; - } - for (int index = 0; index < SUMS_PER_SUPER_BLOCK; index++) { - sums[blockSumOffset + index] = 0; - } - scales[rowScaleOffset + block] = 0.0f; - continue; - } - float inverseScale = -127.0f / extremum; - int partialSum = 0; - for (int index = 0; index < SUPER_BLOCK_VALUES; index++) { - int quant = ggmlNearestInt(inverseScale * input[offset + index]); - byte stored = (byte) Math.min(127, quant); - activations[offset + index] = stored; - partialSum += stored; - if ((index + 1) % SUM_BLOCK_VALUES == 0) { - sums[blockSumOffset + index / SUM_BLOCK_VALUES] = partialSum; - partialSum = 0; - } - } - scales[rowScaleOffset + block] = 1.0f / inverseScale; - } - } - } - - /** Validates one Q4_K projection's device storage before a task graph is built. */ - public static void validateQ4K( - ByteArray weights, - ByteArray activations, - FloatArray activationScales, - IntArray activationSums, - FloatArray output, - int batchSize, - int rows, - int cols) { - validateShape(batchSize, rows, cols); - validateWeights(weights, rows, cols, Q4_K_BLOCK_BYTES); - validateActivations(activations, activationScales, batchSize, cols); - if (activationSums == null - || activationSums.getSize() != Math.multiplyExact(batchSize, cols / SUM_BLOCK_VALUES)) { - throw new IllegalArgumentException("activation sums do not match the projection shape"); - } - validateOutput(output, batchSize, rows); - } - - /** Validates one Q6_K projection's device storage before a task graph is built. */ - public static void validateQ6K( - ByteArray weights, - ByteArray activations, - FloatArray activationScales, - FloatArray output, - int batchSize, - int rows, - int cols) { - validateShape(batchSize, rows, cols); - validateWeights(weights, rows, cols, Q6_K_BLOCK_BYTES); - validateActivations(activations, activationScales, batchSize, cols); - validateOutput(output, batchSize, rows); - } - - private static void validateShape(int batchSize, int rows, int cols) { - if (batchSize < 1) { - throw new IllegalArgumentException("batchSize must be positive"); - } - if (rows < 1) { - throw new IllegalArgumentException("rows must be positive"); - } - if (cols < SUPER_BLOCK_VALUES || cols % SUPER_BLOCK_VALUES != 0) { - throw new IllegalArgumentException("cols must be a positive multiple of 256"); - } - } - - private static void validateWeights(ByteArray weights, int rows, int cols, int blockBytes) { - if (weights == null) { - throw new IllegalArgumentException("weights do not match the projection shape"); - } - int blocks = cols / SUPER_BLOCK_VALUES; - long expected = (long) rows * blocks * blockBytes; - if (expected > Integer.MAX_VALUE || weights.getSize() != (int) expected) { - throw new IllegalArgumentException("weights do not match the projection shape"); - } - } - - private static void validateActivations( - ByteArray activations, FloatArray activationScales, int batchSize, int cols) { - if (activations == null || activations.getSize() != Math.multiplyExact(batchSize, cols)) { - throw new IllegalArgumentException("activations do not match the projection shape"); - } - if (activationScales == null - || activationScales.getSize() != Math.multiplyExact(batchSize, cols / SUPER_BLOCK_VALUES)) { - throw new IllegalArgumentException("activation scales do not match the projection shape"); - } - } - - private static void validateOutput(FloatArray output, int batchSize, int rows) { - if (output == null || output.getSize() != Math.multiplyExact(batchSize, rows)) { - throw new IllegalArgumentException("output does not match the projection shape"); - } - } - - private static int ggmlNearestInt(float value) { - int bits = Float.floatToRawIntBits(value + 12_582_912.0f); - return (bits & 0x007F_FFFF) - 0x0040_0000; - } -} diff --git a/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/PlanShapeStrategy.java b/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/PlanShapeStrategy.java deleted file mode 100644 index 9fa2a34f..00000000 --- a/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/PlanShapeStrategy.java +++ /dev/null @@ -1,108 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.backend.tornado; - -/** - * How retained TornadoVM execution plans hold model weights on the device. - * - *

Only {@link #PER_SHAPE_WHOLE_MODEL} is implemented today. The other constants exist so the - * capacity gate can say precisely which unimplemented plan shape a model would have needed, instead - * of reporting an anonymous shortfall. - */ -public enum PlanShapeStrategy { - - /** - * One {@code TaskGraph} per (weight tensor, batch shape), each owning its own device weight - * buffer and its own host copy. - * - *

This is what {@code TornadoGgufBatchedMatrixKernel} builds: the plan cache is keyed on the - * weight address plus the execution batch size, and the constructor calls {@code - * ByteArray.fromSegment(weights)}, which copies the tensor into a fresh off-heap array before the - * task graph pins it with {@code DataTransferMode.FIRST_EXECUTION}. With decode acceleration on, - * every weight is therefore held twice on the device and twice on the host. - */ - PER_SHAPE_WHOLE_MODEL(true, null), - - /** - * One device weight buffer per tensor, shared by every batch shape that reads it. - * - *

Halves both device and host weight residency when prefill and decode plans are retained - * together. - */ - SHARED_WEIGHT_UPLOAD( - false, - "TornadoGgufBatchedMatrixKernel allocates a new ByteArray per (weight, batch shape) plan;" - + " sharing one device buffer across shapes requires hoisting the weight array out of" - + " the plan constructor and keying it on the weight address alone"), - - /** - * A bounded resident working set: weights stay on the host and only the tensors needed by the - * current layer or routed experts are uploaded, under an LRU bound. - * - *

Device capacity stops being the binding constraint and interconnect bandwidth becomes it, so - * this strategy is never selected by the capacity gate; it is named only so an ineligibility - * message can point at it. - */ - BOUNDED_RESIDENT_WORKING_SET( - false, - "no eviction path exists: plans are held in unbounded LinkedHashMaps and only released on" - + " kernel close, and per-token re-upload is bounded by PCIe bandwidth rather than by" - + " device capacity"); - - private final boolean implemented; - private final String limitation; - - PlanShapeStrategy(boolean implemented, String limitation) { - this.implemented = implemented; - this.limitation = limitation; - } - - /** Whether the shipped kernel can actually build plans in this shape. */ - public boolean implemented() { - return implemented; - } - - /** What stops this strategy from being available, or {@code null} when it is implemented. */ - public String limitation() { - return limitation; - } - - /** Whether device capacity for this strategy can be computed from a request. */ - public boolean capacityComputable() { - return this != BOUNDED_RESIDENT_WORKING_SET; - } - - /** Device bytes of model weights retained under this strategy. */ - long deviceWeightBytes(DeviceMemoryRequest request) { - return switch (this) { - case PER_SHAPE_WHOLE_MODEL -> - Math.multiplyExact(request.weightBytes(), request.retainedShapes()); - case SHARED_WEIGHT_UPLOAD -> request.weightBytes(); - case BOUNDED_RESIDENT_WORKING_SET -> - throw new IllegalStateException( - "BOUNDED_RESIDENT_WORKING_SET has no capacity-derived weight budget"); - }; - } - - /** - * Host bytes of off-heap weight copies retained under this strategy. - * - *

{@code ByteArray.fromSegment} copies; the mapped GGUF stays mapped alongside these copies. - */ - long hostWeightCopyBytes(DeviceMemoryRequest request) { - return deviceWeightBytes(request); - } -} diff --git a/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/Q4ProjectionKernel.java b/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/Q4ProjectionKernel.java deleted file mode 100644 index 22d37986..00000000 --- a/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/Q4ProjectionKernel.java +++ /dev/null @@ -1,348 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.backend.tornado; - -import uk.ac.manchester.tornado.api.annotations.Parallel; -import uk.ac.manchester.tornado.api.types.arrays.ByteArray; -import uk.ac.manchester.tornado.api.types.arrays.FloatArray; - -/** Java-authored TornadoVM kernel for production-compatible Q4_0 by Q8_0 projections. */ -public final class Q4ProjectionKernel { - private static final int BLOCK_VALUES = 32; - private static final int BLOCK_BYTES = 18; - - private Q4ProjectionKernel() {} - - /** Runs one work item per batch/output-row pair over host-prepared Q8_0 activations. */ - public static void multiply( - ByteArray weights, - ByteArray activations, - FloatArray activationScales, - FloatArray output, - int batchSize, - int rows, - int cols) { - int blocks = cols / BLOCK_VALUES; - for (@Parallel int outputIndex = 0; outputIndex < batchSize * rows; outputIndex++) { - int batch = outputIndex / rows; - int row = outputIndex - batch * rows; - int activationOffset = batch * cols; - int activationScaleOffset = batch * blocks; - output.set( - outputIndex, - rowDotLow( - weights, - row, - activations, - activationOffset, - activationScales, - activationScaleOffset, - blocks) - + rowDotHigh( - weights, - row, - activations, - activationOffset, - activationScales, - activationScaleOffset, - blocks)); - } - } - - /** Computes two projections in one dispatch over one prepared activation. */ - public static void multiplyDual( - ByteArray firstWeights, - int firstRows, - ByteArray secondWeights, - int secondRows, - ByteArray activations, - FloatArray activationScales, - FloatArray firstOutput, - FloatArray secondOutput, - int batchSize, - int cols) { - int blocks = cols / BLOCK_VALUES; - int combinedRows = firstRows + secondRows; - for (@Parallel int outputIndex = 0; outputIndex < batchSize * combinedRows; outputIndex++) { - int batch = outputIndex / combinedRows; - int combinedRow = outputIndex - batch * combinedRows; - int activationOffset = batch * cols; - int activationScaleOffset = batch * blocks; - if (combinedRow < firstRows) { - firstOutput.set( - batch * firstRows + combinedRow, - rowDotLow( - firstWeights, - combinedRow, - activations, - activationOffset, - activationScales, - activationScaleOffset, - blocks) - + rowDotHigh( - firstWeights, - combinedRow, - activations, - activationOffset, - activationScales, - activationScaleOffset, - blocks)); - } else { - int row = combinedRow - firstRows; - secondOutput.set( - batch * secondRows + row, - rowDotLow( - secondWeights, - row, - activations, - activationOffset, - activationScales, - activationScaleOffset, - blocks) - + rowDotHigh( - secondWeights, - row, - activations, - activationOffset, - activationScales, - activationScaleOffset, - blocks)); - } - } - } - - /** Computes three projections in one dispatch over one prepared activation. */ - public static void multiplyTriple( - ByteArray firstWeights, - int firstRows, - ByteArray secondWeights, - int secondRows, - ByteArray thirdWeights, - int thirdRows, - ByteArray activations, - FloatArray activationScales, - FloatArray firstOutput, - FloatArray secondOutput, - FloatArray thirdOutput, - int batchSize, - int cols) { - int blocks = cols / BLOCK_VALUES; - int firstAndSecondRows = firstRows + secondRows; - int combinedRows = firstAndSecondRows + thirdRows; - for (@Parallel int outputIndex = 0; outputIndex < batchSize * combinedRows; outputIndex++) { - int batch = outputIndex / combinedRows; - int combinedRow = outputIndex - batch * combinedRows; - int activationOffset = batch * cols; - int activationScaleOffset = batch * blocks; - if (combinedRow < firstRows) { - firstOutput.set( - batch * firstRows + combinedRow, - rowDotLow( - firstWeights, - combinedRow, - activations, - activationOffset, - activationScales, - activationScaleOffset, - blocks) - + rowDotHigh( - firstWeights, - combinedRow, - activations, - activationOffset, - activationScales, - activationScaleOffset, - blocks)); - } else if (combinedRow < firstAndSecondRows) { - int row = combinedRow - firstRows; - secondOutput.set( - batch * secondRows + row, - rowDotLow( - secondWeights, - row, - activations, - activationOffset, - activationScales, - activationScaleOffset, - blocks) - + rowDotHigh( - secondWeights, - row, - activations, - activationOffset, - activationScales, - activationScaleOffset, - blocks)); - } else { - int row = combinedRow - firstAndSecondRows; - thirdOutput.set( - batch * thirdRows + row, - rowDotLow( - thirdWeights, - row, - activations, - activationOffset, - activationScales, - activationScaleOffset, - blocks) - + rowDotHigh( - thirdWeights, - row, - activations, - activationOffset, - activationScales, - activationScaleOffset, - blocks)); - } - } - } - - private static float rowDotLow( - ByteArray weights, - int row, - ByteArray activations, - int activationOffset, - FloatArray activationScales, - int activationScaleOffset, - int blocks) { - int rowBlockOffset = row * blocks; - float sum = 0.0f; - for (int block = 0; block < blocks; block++) { - int weightOffset = (rowBlockOffset + block) * BLOCK_BYTES; - int quantOffset = activationOffset + block * BLOCK_VALUES; - int integerSum = pairProductSumLow(weights, weightOffset, activations, quantOffset); - float scale = - weights.getHalfFloat(weightOffset).getFloat32() - * activationScales.get(activationScaleOffset + block); - sum += scale * integerSum; - } - return sum; - } - - private static float rowDotHigh( - ByteArray weights, - int row, - ByteArray activations, - int activationOffset, - FloatArray activationScales, - int activationScaleOffset, - int blocks) { - int rowBlockOffset = row * blocks; - float sum = 0.0f; - for (int block = 0; block < blocks; block++) { - int weightOffset = (rowBlockOffset + block) * BLOCK_BYTES; - int quantOffset = activationOffset + block * BLOCK_VALUES; - int integerSum = pairProductSumHigh(weights, weightOffset, activations, quantOffset); - float scale = - weights.getHalfFloat(weightOffset).getFloat32() - * activationScales.get(activationScaleOffset + block); - sum += scale * integerSum; - } - return sum; - } - - private static int pairProductSumLow( - ByteArray weights, int weightOffset, ByteArray activations, int activationOffset) { - return pairProduct(weights, weightOffset, activations, activationOffset, 0) - + pairProduct(weights, weightOffset, activations, activationOffset, 1) - + pairProduct(weights, weightOffset, activations, activationOffset, 2) - + pairProduct(weights, weightOffset, activations, activationOffset, 3) - + pairProduct(weights, weightOffset, activations, activationOffset, 4) - + pairProduct(weights, weightOffset, activations, activationOffset, 5) - + pairProduct(weights, weightOffset, activations, activationOffset, 6) - + pairProduct(weights, weightOffset, activations, activationOffset, 7); - } - - private static int pairProductSumHigh( - ByteArray weights, int weightOffset, ByteArray activations, int activationOffset) { - return pairProduct(weights, weightOffset, activations, activationOffset, 8) - + pairProduct(weights, weightOffset, activations, activationOffset, 9) - + pairProduct(weights, weightOffset, activations, activationOffset, 10) - + pairProduct(weights, weightOffset, activations, activationOffset, 11) - + pairProduct(weights, weightOffset, activations, activationOffset, 12) - + pairProduct(weights, weightOffset, activations, activationOffset, 13) - + pairProduct(weights, weightOffset, activations, activationOffset, 14) - + pairProduct(weights, weightOffset, activations, activationOffset, 15); - } - - private static int pairProduct( - ByteArray weights, int weightOffset, ByteArray activations, int activationOffset, int index) { - int packed = weights.get(weightOffset + 2 + index) & 0xFF; - int low = (packed & 0x0F) - 8; - int high = ((packed >>> 4) & 0x0F) - 8; - return low * activations.get(activationOffset + index) - + high * activations.get(activationOffset + index + 16); - } - - /** Reproduces the production Q8_0 activation preparation, including FP16 scale rounding. */ - public static void quantize( - float[] input, byte[] activations, float[] scales, int batchSize, int cols) { - int blocks = cols / BLOCK_VALUES; - for (int batch = 0; batch < batchSize; batch++) { - for (int block = 0; block < blocks; block++) { - int offset = batch * cols + block * BLOCK_VALUES; - float absoluteMax = 0.0f; - for (int index = 0; index < BLOCK_VALUES; index++) { - absoluteMax = Math.max(absoluteMax, Math.abs(input[offset + index])); - } - float scale = absoluteMax / 127.0f; - float inverseScale = absoluteMax == 0.0f ? 0.0f : 127.0f / absoluteMax; - scales[batch * blocks + block] = Float.float16ToFloat(Float.floatToFloat16(scale)); - for (int index = 0; index < BLOCK_VALUES; index++) { - activations[offset + index] = (byte) ggmlNearestInt(input[offset + index] * inverseScale); - } - } - } - } - - public static void validate( - ByteArray weights, - ByteArray activations, - FloatArray activationScales, - FloatArray output, - int batchSize, - int rows, - int cols) { - if (batchSize < 1) { - throw new IllegalArgumentException("batchSize must be positive"); - } - if (rows < 1) { - throw new IllegalArgumentException("rows must be positive"); - } - if (cols < BLOCK_VALUES || cols % BLOCK_VALUES != 0) { - throw new IllegalArgumentException("cols must be a positive multiple of 32"); - } - int blocks = cols / BLOCK_VALUES; - int expectedWeightBytes = Math.multiplyExact(Math.multiplyExact(rows, blocks), BLOCK_BYTES); - if (weights.getSize() != expectedWeightBytes) { - throw new IllegalArgumentException("weights do not match the projection shape"); - } - if (activations.getSize() != Math.multiplyExact(batchSize, cols)) { - throw new IllegalArgumentException("activations do not match the projection shape"); - } - if (activationScales.getSize() != Math.multiplyExact(batchSize, blocks)) { - throw new IllegalArgumentException("activation scales do not match the projection shape"); - } - if (output.getSize() != Math.multiplyExact(batchSize, rows)) { - throw new IllegalArgumentException("output does not match the projection shape"); - } - } - - private static int ggmlNearestInt(float value) { - int bits = Float.floatToRawIntBits(value + 12_582_912.0f); - return (bits & 0x007F_FFFF) - 0x0040_0000; - } -} diff --git a/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/TornadoBackend.java b/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/TornadoBackend.java deleted file mode 100644 index 25039add..00000000 --- a/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/TornadoBackend.java +++ /dev/null @@ -1,172 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.backend.tornado; - -import com.integrallis.models.api.BackendConfiguration; -import com.integrallis.models.backend.purejava.PureJavaBackend; -import com.integrallis.models.backend.purejava.spi.GgufBatchedMatrixKernel; -import java.io.IOException; -import java.io.UncheckedIOException; -import java.nio.file.Files; -import java.nio.file.Path; -import java.time.Duration; -import java.util.Arrays; -import java.util.List; -import java.util.Objects; - -/** Loads the qualified Java/Tornado Q4 backend when possible and otherwise uses the Vector API. */ -public final class TornadoBackend { - private static final System.Logger LOGGER = System.getLogger(TornadoBackend.class.getName()); - - private TornadoBackend() {} - - /** Opens a model with automatic device selection, eager readiness, and safe CPU fallback. */ - public static TornadoBackendRuntime open(Path modelPath) { - return open(modelPath, BackendConfiguration.empty(), TornadoBackendOptions.defaults()); - } - - /** Opens a model using explicit accelerator readiness and fallback controls. */ - public static TornadoBackendRuntime open(Path modelPath, TornadoBackendOptions options) { - return open(modelPath, BackendConfiguration.empty(), options); - } - - /** Opens a model using its qualified backend configuration and accelerator controls. */ - public static TornadoBackendRuntime open( - Path modelPath, BackendConfiguration backendConfiguration, TornadoBackendOptions options) { - Objects.requireNonNull(modelPath, "modelPath"); - Objects.requireNonNull(backendConfiguration, "backendConfiguration"); - Objects.requireNonNull(options, "options"); - Path model = modelPath.toAbsolutePath().normalize(); - if (!Files.isRegularFile(model)) { - throw new IllegalArgumentException("modelPath must be a regular model file: " + model); - } - long modelBytes = size(model); - List devices; - try { - devices = TornadoRuntimeDevices.discover(); - } catch (LinkageError | RuntimeException failure) { - return fallback( - model, - backendConfiguration, - options, - "TornadoVM runtime unavailable: " + failure.getClass().getSimpleName(), - 0); - } - AcceleratorEligibility.Decision decision = - AcceleratorEligibility.select( - devices, - DeviceMemoryRequest.ofModelFile( - modelLabel(model), modelBytes, options.accelerateDecode())); - if (!decision.eligible()) { - return fallback( - model, backendConfiguration, options, decision.reason(), decision.requiredBytes()); - } - - TornadoGgufBatchedMatrixKernel kernel = - new TornadoGgufBatchedMatrixKernel( - options.executionBatchSize(), options.accelerateDecode()); - PureJavaBackend backend = null; - try { - backend = PureJavaBackend.load(model, backendConfiguration, kernel); - Duration readiness = options.eagerReadiness() ? prepare(backend, options) : Duration.ZERO; - if (options.eagerReadiness() && kernel.calls() == 0) { - backend.close(); - return fallback( - model, - backendConfiguration, - options, - "model has no eligible Q4_0 or K-quant projections", - decision.requiredBytes()); - } - String device = decision.device().name(); - LOGGER.log( - System.Logger.Level.INFO, - "Models accelerator selected {0}; readiness={1} ms plans={2} routed={3}", - device, - readiness.toMillis(), - kernel.projectionPlanCount(), - kernel.routedProjectionsByFormat()); - return new TornadoBackendRuntime( - backend, - new TornadoBackendStatus(true, device, "eligible", decision.requiredBytes(), readiness), - kernel); - } catch (LinkageError | RuntimeException failure) { - if (backend != null) { - try { - backend.close(); - } catch (RuntimeException closeFailure) { - failure.addSuppressed(closeFailure); - } - } - if (options.requireAccelerator()) { - throw new IllegalStateException("qualified accelerator initialization failed", failure); - } - return fallback( - model, - backendConfiguration, - options, - "accelerator initialization failed: " + failure.getClass().getSimpleName(), - decision.requiredBytes()); - } - } - - private static Duration prepare(PureJavaBackend backend, TornadoBackendOptions options) { - long started = System.nanoTime(); - int token = readinessToken(backend); - int[] tokens = new int[options.executionBatchSize()]; - Arrays.fill(tokens, token); - backend.prefill(tokens, 0); - if (options.accelerateDecode()) { - backend.forward(token, tokens.length); - } - backend.reset(); - return Duration.ofNanos(System.nanoTime() - started); - } - - private static int readinessToken(PureJavaBackend backend) { - int token = backend.tokenizer().bosToken(); - return token >= 0 && token < backend.metadata().vocabSize() ? token : 0; - } - - private static TornadoBackendRuntime fallback( - Path model, - BackendConfiguration backendConfiguration, - TornadoBackendOptions options, - String reason, - long requiredBytes) { - if (options.requireAccelerator()) { - throw new IllegalStateException("accelerator required but unavailable: " + reason); - } - LOGGER.log( - System.Logger.Level.INFO, "Models accelerator unavailable; using Vector API: {0}", reason); - return new TornadoBackendRuntime( - PureJavaBackend.load(model, backendConfiguration, GgufBatchedMatrixKernel.none()), - new TornadoBackendStatus(false, "Vector API", reason, requiredBytes, Duration.ZERO)); - } - - private static String modelLabel(Path model) { - Path fileName = model.getFileName(); - return fileName == null ? model.toString() : fileName.toString(); - } - - private static long size(Path model) { - try { - return Files.size(model); - } catch (IOException exception) { - throw new UncheckedIOException("could not read model size: " + model, exception); - } - } -} diff --git a/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/TornadoBackendOptions.java b/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/TornadoBackendOptions.java deleted file mode 100644 index eab6fbd5..00000000 --- a/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/TornadoBackendOptions.java +++ /dev/null @@ -1,35 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.backend.tornado; - -/** Immutable controls for optional TornadoVM device selection and readiness. */ -public record TornadoBackendOptions( - boolean accelerateDecode, - boolean eagerReadiness, - boolean requireAccelerator, - int executionBatchSize) { - - public TornadoBackendOptions { - if (executionBatchSize < 4) { - throw new IllegalArgumentException("executionBatchSize must be at least 4"); - } - } - - /** Selects qualified hardware automatically and falls back to the Java Vector API. */ - public static TornadoBackendOptions defaults() { - return new TornadoBackendOptions(true, true, false, 32); - } -} diff --git a/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/TornadoBackendRuntime.java b/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/TornadoBackendRuntime.java deleted file mode 100644 index b0bc7eb4..00000000 --- a/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/TornadoBackendRuntime.java +++ /dev/null @@ -1,81 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.backend.tornado; - -import com.integrallis.models.backend.purejava.PureJavaBackend; -import java.util.Map; -import java.util.Objects; - -/** Owns the automatically selected accelerated or Vector API backend. */ -public final class TornadoBackendRuntime implements AutoCloseable { - private PureJavaBackend backend; - private final TornadoBackendStatus status; - private final TornadoGgufBatchedMatrixKernel kernel; - - TornadoBackendRuntime(PureJavaBackend backend, TornadoBackendStatus status) { - this(backend, status, null); - } - - TornadoBackendRuntime( - PureJavaBackend backend, TornadoBackendStatus status, TornadoGgufBatchedMatrixKernel kernel) { - this.backend = Objects.requireNonNull(backend, "backend"); - this.status = Objects.requireNonNull(status, "status"); - this.kernel = kernel; - } - - /** Returns the loaded backend used by the ordinary Models generation pipeline. */ - public PureJavaBackend backend() { - if (backend == null) { - throw new IllegalStateException("backend ownership was transferred"); - } - return backend; - } - - PureJavaBackend detachBackend() { - PureJavaBackend detached = backend(); - backend = null; - return detached; - } - - /** Returns the device-selection, fallback, and readiness outcome. */ - public TornadoBackendStatus status() { - return status; - } - - /** - * Returns the projections routed to the device so far, keyed by GGUF weight format. - * - *

Empty when the load fell back to the Vector API. A grouped dispatch counts once per matrix, - * so a mixed-format model's counts show whether each of its formats reached the device rather - * than only the majority one. - */ - public Map routedProjectionsByFormat() { - return kernel == null ? Map.of() : kernel.routedProjectionsByFormat(); - } - - /** Number of distinct compiled device execution plans, or zero when not accelerated. */ - public int projectionPlanCount() { - return kernel == null ? 0 : kernel.projectionPlanCount(); - } - - @Override - public void close() { - if (backend != null) { - backend.close(); - backend = null; - } - } -} diff --git a/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/TornadoBackendStatus.java b/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/TornadoBackendStatus.java deleted file mode 100644 index 6b9b260a..00000000 --- a/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/TornadoBackendStatus.java +++ /dev/null @@ -1,48 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.backend.tornado; - -import java.time.Duration; -import java.util.Objects; - -/** Immutable outcome of automatic accelerator selection and readiness. */ -public record TornadoBackendStatus( - boolean accelerated, - String device, - String reason, - long requiredDeviceBytes, - Duration readinessTime) { - - public TornadoBackendStatus { - device = requireText(device, "device"); - reason = requireText(reason, "reason"); - if (requiredDeviceBytes < 0) { - throw new IllegalArgumentException("requiredDeviceBytes must not be negative"); - } - readinessTime = Objects.requireNonNull(readinessTime, "readinessTime"); - if (readinessTime.isNegative()) { - throw new IllegalArgumentException("readinessTime must not be negative"); - } - } - - private static String requireText(String value, String name) { - Objects.requireNonNull(value, name); - if (value.isBlank()) { - throw new IllegalArgumentException(name + " must not be blank"); - } - return value.strip(); - } -} diff --git a/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/TornadoGgufBatchedMatrixKernel.java b/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/TornadoGgufBatchedMatrixKernel.java deleted file mode 100644 index d4d27235..00000000 --- a/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/TornadoGgufBatchedMatrixKernel.java +++ /dev/null @@ -1,1091 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.backend.tornado; - -import com.integrallis.models.backend.purejava.gguf.GgufTensorType; -import com.integrallis.models.backend.purejava.plan.PureJavaPlanConfiguration; -import com.integrallis.models.backend.purejava.spi.GgufBatchedMatrixKernel; -import java.lang.foreign.MemorySegment; -import java.util.ArrayList; -import java.util.Arrays; -import java.util.Collections; -import java.util.EnumMap; -import java.util.LinkedHashMap; -import java.util.List; -import java.util.Map; -import java.util.Objects; -import uk.ac.manchester.tornado.api.TaskGraph; -import uk.ac.manchester.tornado.api.TornadoExecutionPlan; -import uk.ac.manchester.tornado.api.enums.DataTransferMode; -import uk.ac.manchester.tornado.api.types.arrays.ByteArray; -import uk.ac.manchester.tornado.api.types.arrays.FloatArray; -import uk.ac.manchester.tornado.api.types.arrays.IntArray; - -/** - * Q4_0 and K-quant prefill and optional decode projections backed by Java-authored TornadoVM - * kernels. - * - *

Two activation families are served. Q4_0 weights consume the Q8_0 activations prepared by - * {@link Q4ProjectionKernel}; Q4_K and Q6_K weights consume the Q8_K activations prepared by {@link - * KQuantProjectionKernel}. A grouped dispatch shares one activation preparation, so the two - * families are never mixed inside one dual or triple projection even though a Q4_K_M model contains - * both. - */ -public final class TornadoGgufBatchedMatrixKernel implements GgufBatchedMatrixKernel { - private static final int MINIMUM_BATCH = 4; - private static final int DEFAULT_EXECUTION_BATCH_SIZE = 32; - private static final long MINIMUM_MATRIX_VALUES = 1_048_576L; - - /** - * Largest weight tensor a single execution plan can hold on the device. - * - *

TornadoVM's {@code ByteArray} is indexed with an {@code int} and carries a small header, so - * a mapped tensor at or above two gibibytes cannot be addressed at all. Rejecting it here keeps - * the projection on the Vector API rather than failing the load. - */ - private static final long MAX_DEVICE_TENSOR_BYTES = Integer.MAX_VALUE - 1024L; - - private static final Map PLAN_RECOMMENDATIONS = - Map.of( - PureJavaPlanConfiguration.GROUPED_PROJECTIONS_PROPERTY, - "true", - PureJavaPlanConfiguration.MIXED_K_PROJECTIONS_PROPERTY, - "true", - PureJavaPlanConfiguration.STAGED_QUANTIZED_FFN_PROPERTY, - "false", - PureJavaPlanConfiguration.STAGED_QUANTIZED_LAYER_PROPERTY, - "false"); - - private final Map plans = new LinkedHashMap<>(); - private final Map dualPlans = new LinkedHashMap<>(); - private final Map triplePlans = new LinkedHashMap<>(); - private final Map kQuantPlans = new LinkedHashMap<>(); - private final Map routedProjections = new EnumMap<>(GgufTensorType.class); - private final int executionBatchSize; - private final boolean accelerateDecode; - private int planSequence; - private long calls; - private long totalNanos; - private boolean closed; - - public TornadoGgufBatchedMatrixKernel() { - this(DEFAULT_EXECUTION_BATCH_SIZE, false); - } - - public TornadoGgufBatchedMatrixKernel(int executionBatchSize) { - this(executionBatchSize, false); - } - - public TornadoGgufBatchedMatrixKernel(int executionBatchSize, boolean accelerateDecode) { - if (executionBatchSize < MINIMUM_BATCH) { - throw new IllegalArgumentException("executionBatchSize must be at least " + MINIMUM_BATCH); - } - this.executionBatchSize = executionBatchSize; - this.accelerateDecode = accelerateDecode; - } - - /** Fixed device batch shape used to make compiled plans reusable across prompt lengths. */ - public int executionBatchSize() { - return executionBatchSize; - } - - /** Whether this provider creates separate single-token projection plans for decode. */ - public boolean acceleratesDecode() { - return accelerateDecode; - } - - int executionBatchSizeFor(int actualBatchSize) { - return accelerateDecode && actualBatchSize == 1 ? 1 : executionBatchSize; - } - - @Override - public String implementation() { - return accelerateDecode ? "tornadovm-java-q4-prefill-decode" : "tornadovm-java-q4-prefill"; - } - - @Override - public Map planRecommendations() { - return PLAN_RECOMMENDATIONS; - } - - @Override - public boolean supports(GgufTensorType type) { - return type == GgufTensorType.Q4_0 - || type == GgufTensorType.Q4_K - || type == GgufTensorType.Q6_K; - } - - @Override - public boolean isEligible(GgufTensorType type, int batchSize, int rows, int cols) { - return supports(type) - && eligibleBatchSize(batchSize) - && batchSize <= executionBatchSize - && rows > 0 - && cols > 0 - && cols % type.blockSize() == 0 - && deviceAddressable(type, rows, cols) - && (long) rows * cols >= MINIMUM_MATRIX_VALUES; - } - - @Override - public boolean supportsDual(GgufTensorType firstType, GgufTensorType secondType) { - if (firstType == GgufTensorType.Q4_0) { - return secondType == GgufTensorType.Q4_0; - } - return firstType == GgufTensorType.Q4_K && secondType == GgufTensorType.Q4_K; - } - - @Override - public boolean isDualEligible( - GgufTensorType firstType, - int firstRows, - GgufTensorType secondType, - int secondRows, - int batchSize, - int cols) { - return supportsDual(firstType, secondType) - && cols % firstType.blockSize() == 0 - && deviceAddressable(firstType, firstRows, cols) - && deviceAddressable(secondType, secondRows, cols) - && eligibleCombined(batchSize, cols, firstRows, secondRows); - } - - @Override - public synchronized void multiplyDual( - float[] firstOutput, - MemorySegment firstWeights, - GgufTensorType firstType, - int firstRows, - float[] secondOutput, - MemorySegment secondWeights, - GgufTensorType secondType, - int secondRows, - float[] input, - int batchSize, - int cols) { - requireOpen(); - if (!isDualEligible(firstType, firstRows, secondType, secondRows, batchSize, cols)) { - throw new UnsupportedOperationException( - "dual projection is not eligible for the Tornado backend"); - } - validateProjectionStorage(firstOutput, firstWeights, input, batchSize, firstRows, cols); - validateProjectionStorage(secondOutput, secondWeights, input, batchSize, secondRows, cols); - int planBatchSize = executionBatchSizeFor(batchSize); - long started = System.nanoTime(); - if (isKQuant(firstType)) { - kQuantPlan( - new MemorySegment[] {firstWeights, secondWeights}, - new GgufTensorType[] {firstType, secondType}, - new int[] {firstRows, secondRows}, - planBatchSize, - cols) - .execute(input, new float[][] {firstOutput, secondOutput}, batchSize); - } else { - DualProjectionKey key = - new DualProjectionKey( - firstWeights.address(), - firstWeights.byteSize(), - firstRows, - secondWeights.address(), - secondWeights.byteSize(), - secondRows, - planBatchSize, - cols); - DualProjectionPlan plan = - dualPlans.computeIfAbsent( - key, - ignored -> - new DualProjectionPlan( - nextPlanName(), - firstWeights, - firstRows, - secondWeights, - secondRows, - planBatchSize, - cols)); - plan.execute(input, firstOutput, secondOutput, batchSize); - } - recordCall(started, firstType, secondType); - } - - @Override - public boolean supportsTriple( - GgufTensorType firstType, GgufTensorType secondType, GgufTensorType thirdType) { - if (firstType == GgufTensorType.Q4_0) { - return secondType == GgufTensorType.Q4_0 && thirdType == GgufTensorType.Q4_0; - } - if (firstType != GgufTensorType.Q4_K || secondType != GgufTensorType.Q4_K) { - return false; - } - // Q4_K_M promotes the value projection to Q6_K while query and key stay Q4_K. - return thirdType == GgufTensorType.Q4_K || thirdType == GgufTensorType.Q6_K; - } - - @Override - public boolean isTripleEligible( - GgufTensorType firstType, - int firstRows, - GgufTensorType secondType, - int secondRows, - GgufTensorType thirdType, - int thirdRows, - int batchSize, - int cols) { - return supportsTriple(firstType, secondType, thirdType) - && cols % firstType.blockSize() == 0 - && cols % thirdType.blockSize() == 0 - && deviceAddressable(firstType, firstRows, cols) - && deviceAddressable(secondType, secondRows, cols) - && deviceAddressable(thirdType, thirdRows, cols) - && eligibleCombined(batchSize, cols, firstRows, secondRows, thirdRows); - } - - @Override - public synchronized void multiplyTriple( - float[] firstOutput, - MemorySegment firstWeights, - GgufTensorType firstType, - int firstRows, - float[] secondOutput, - MemorySegment secondWeights, - GgufTensorType secondType, - int secondRows, - float[] thirdOutput, - MemorySegment thirdWeights, - GgufTensorType thirdType, - int thirdRows, - float[] input, - int batchSize, - int cols) { - requireOpen(); - if (!isTripleEligible( - firstType, firstRows, secondType, secondRows, thirdType, thirdRows, batchSize, cols)) { - throw new UnsupportedOperationException( - "triple projection is not eligible for the Tornado backend"); - } - validateProjectionStorage(firstOutput, firstWeights, input, batchSize, firstRows, cols); - validateProjectionStorage(secondOutput, secondWeights, input, batchSize, secondRows, cols); - validateProjectionStorage(thirdOutput, thirdWeights, input, batchSize, thirdRows, cols); - int planBatchSize = executionBatchSizeFor(batchSize); - long started = System.nanoTime(); - if (isKQuant(firstType)) { - kQuantPlan( - new MemorySegment[] {firstWeights, secondWeights, thirdWeights}, - new GgufTensorType[] {firstType, secondType, thirdType}, - new int[] {firstRows, secondRows, thirdRows}, - planBatchSize, - cols) - .execute(input, new float[][] {firstOutput, secondOutput, thirdOutput}, batchSize); - } else { - TripleProjectionKey key = - new TripleProjectionKey( - firstWeights.address(), - firstWeights.byteSize(), - firstRows, - secondWeights.address(), - secondWeights.byteSize(), - secondRows, - thirdWeights.address(), - thirdWeights.byteSize(), - thirdRows, - planBatchSize, - cols); - TripleProjectionPlan plan = - triplePlans.computeIfAbsent( - key, - ignored -> - new TripleProjectionPlan( - nextPlanName(), - firstWeights, - firstRows, - secondWeights, - secondRows, - thirdWeights, - thirdRows, - planBatchSize, - cols)); - plan.execute(input, firstOutput, secondOutput, thirdOutput, batchSize); - } - recordCall(started, firstType, secondType, thirdType); - } - - @Override - public synchronized void multiply( - float[] output, - float[] input, - MemorySegment weights, - GgufTensorType type, - int batchSize, - int rows, - int cols) { - requireOpen(); - if (!isEligible(type, batchSize, rows, cols)) { - throw new UnsupportedOperationException("projection is not eligible for the Tornado backend"); - } - validateProjectionStorage(output, weights, input, batchSize, rows, cols); - int planBatchSize = executionBatchSizeFor(batchSize); - long started = System.nanoTime(); - if (isKQuant(type)) { - kQuantPlan( - new MemorySegment[] {weights}, - new GgufTensorType[] {type}, - new int[] {rows}, - planBatchSize, - cols) - .execute(input, new float[][] {output}, batchSize); - } else { - ProjectionKey key = - new ProjectionKey(weights.address(), weights.byteSize(), planBatchSize, rows, cols); - ProjectionPlan plan = - plans.computeIfAbsent( - key, - ignored -> new ProjectionPlan(nextPlanName(), weights, planBatchSize, rows, cols)); - plan.execute(input, output, batchSize); - } - recordCall(started, type); - } - - /** Number of distinct tensor/shape execution plans compiled or awaiting first compilation. */ - public synchronized int projectionPlanCount() { - return plans.size() + dualPlans.size() + triplePlans.size() + kQuantPlans.size(); - } - - /** - * Number of model projections routed through this provider, by GGUF weight format. - * - *

A grouped dispatch counts once per matrix, so a Q4_K/Q4_K/Q6_K attention group adds two to - * {@code Q4_K} and one to {@code Q6_K}. Formats absent from the map were never accelerated, which - * is how a run shows that a mixed-format model actually took the device path for each of its - * formats rather than silently falling back for one of them. - */ - public synchronized Map routedProjectionsByFormat() { - Map byFormat = new LinkedHashMap<>(); - routedProjections.forEach((type, count) -> byFormat.put(type.name(), count)); - return Collections.unmodifiableMap(byFormat); - } - - /** Number of model projection calls routed through this provider. */ - public synchronized long calls() { - return calls; - } - - /** Wall-clock time spent preparing, compiling, executing, and copying accelerated calls. */ - public synchronized double totalMillis() { - return totalNanos / 1_000_000.0; - } - - @Override - public synchronized void close() { - if (closed) { - return; - } - RuntimeException failure = null; - for (ProjectionPlan plan : plans.values()) { - try { - plan.close(); - } catch (RuntimeException exception) { - if (failure == null) { - failure = exception; - } else { - failure.addSuppressed(exception); - } - } - } - for (DualProjectionPlan plan : dualPlans.values()) { - try { - plan.close(); - } catch (RuntimeException exception) { - if (failure == null) { - failure = exception; - } else { - failure.addSuppressed(exception); - } - } - } - for (TripleProjectionPlan plan : triplePlans.values()) { - try { - plan.close(); - } catch (RuntimeException exception) { - if (failure == null) { - failure = exception; - } else { - failure.addSuppressed(exception); - } - } - } - for (KQuantProjectionPlan plan : kQuantPlans.values()) { - try { - plan.close(); - } catch (RuntimeException exception) { - if (failure == null) { - failure = exception; - } else { - failure.addSuppressed(exception); - } - } - } - plans.clear(); - dualPlans.clear(); - triplePlans.clear(); - kQuantPlans.clear(); - closed = true; - if (failure != null) { - throw failure; - } - } - - private void requireOpen() { - if (closed) { - throw new IllegalStateException("Tornado projection kernel is closed"); - } - } - - private boolean eligibleCombined(int batchSize, int cols, int... rows) { - if (!eligibleBatchSize(batchSize) || batchSize > executionBatchSize || cols <= 0) { - return false; - } - long combinedRows = 0; - for (int rowCount : rows) { - if (rowCount <= 0) { - return false; - } - combinedRows += rowCount; - } - return combinedRows * cols >= MINIMUM_MATRIX_VALUES; - } - - private boolean eligibleBatchSize(int batchSize) { - return batchSize >= MINIMUM_BATCH || (accelerateDecode && batchSize == 1); - } - - private static void validateProjectionStorage( - float[] output, MemorySegment weights, float[] input, int batchSize, int rows, int cols) { - Objects.requireNonNull(output, "output"); - Objects.requireNonNull(input, "input"); - Objects.requireNonNull(weights, "weights"); - int expectedInput = Math.multiplyExact(batchSize, cols); - int expectedOutput = Math.multiplyExact(batchSize, rows); - if (input.length < expectedInput || output.length < expectedOutput) { - throw new IllegalArgumentException("input or output storage does not match projection shape"); - } - } - - private String nextPlanName() { - return "q4-model-" + planSequence++; - } - - private void recordCall(long started, GgufTensorType... types) { - totalNanos += System.nanoTime() - started; - calls++; - for (GgufTensorType type : types) { - routedProjections.merge(type, 1L, Long::sum); - } - } - - private static boolean isKQuant(GgufTensorType type) { - return type == GgufTensorType.Q4_K || type == GgufTensorType.Q6_K; - } - - private static boolean deviceAddressable(GgufTensorType type, int rows, int cols) { - if (rows <= 0 || cols <= 0 || cols % type.blockSize() != 0) { - return false; - } - return (long) rows * (cols / type.blockSize()) * type.typeSize() <= MAX_DEVICE_TENSOR_BYTES; - } - - private KQuantProjectionPlan kQuantPlan( - MemorySegment[] weights, GgufTensorType[] types, int[] rows, int planBatchSize, int cols) { - List addresses = new ArrayList<>(weights.length); - List byteSizes = new ArrayList<>(weights.length); - List rowCounts = new ArrayList<>(rows.length); - for (MemorySegment weight : weights) { - addresses.add(weight.address()); - byteSizes.add(weight.byteSize()); - } - for (int rowCount : rows) { - rowCounts.add(rowCount); - } - KQuantProjectionKey key = - new KQuantProjectionKey( - List.copyOf(addresses), - List.copyOf(byteSizes), - List.copyOf(rowCounts), - List.of(types), - planBatchSize, - cols); - return kQuantPlans.computeIfAbsent( - key, - ignored -> - new KQuantProjectionPlan(nextPlanName(), weights, types, rows, planBatchSize, cols)); - } - - private record ProjectionKey(long address, long weightBytes, int batchSize, int rows, int cols) {} - - private record DualProjectionKey( - long firstAddress, - long firstWeightBytes, - int firstRows, - long secondAddress, - long secondWeightBytes, - int secondRows, - int batchSize, - int cols) {} - - private record TripleProjectionKey( - long firstAddress, - long firstWeightBytes, - int firstRows, - long secondAddress, - long secondWeightBytes, - int secondRows, - long thirdAddress, - long thirdWeightBytes, - int thirdRows, - int batchSize, - int cols) {} - - private static final class ProjectionPlan implements AutoCloseable { - private final int batchSize; - private final int rows; - private final int cols; - private final float[] paddedInput; - private final byte[] preparedActivations; - private final float[] preparedScales; - private final ByteArray deviceActivations; - private final FloatArray deviceScales; - private final FloatArray deviceOutput; - private final TornadoExecutionPlan plan; - - private ProjectionPlan(String name, MemorySegment weights, int batchSize, int rows, int cols) { - this.batchSize = batchSize; - this.rows = rows; - this.cols = cols; - int activationEntries = Math.multiplyExact(batchSize, cols); - this.paddedInput = new float[activationEntries]; - this.preparedActivations = new byte[activationEntries]; - this.preparedScales = new float[activationEntries / 32]; - ByteArray deviceWeights = ByteArray.fromSegment(weights); - this.deviceActivations = new ByteArray(activationEntries); - this.deviceScales = new FloatArray(preparedScales.length); - this.deviceOutput = new FloatArray(Math.multiplyExact(batchSize, rows)); - Q4ProjectionKernel.validate( - deviceWeights, deviceActivations, deviceScales, deviceOutput, batchSize, rows, cols); - TaskGraph graph = - new TaskGraph(name) - .transferToDevice(DataTransferMode.FIRST_EXECUTION, deviceWeights) - .transferToDevice(DataTransferMode.EVERY_EXECUTION, deviceActivations, deviceScales) - .task( - "multiply", - Q4ProjectionKernel::multiply, - deviceWeights, - deviceActivations, - deviceScales, - deviceOutput, - batchSize, - rows, - cols) - .transferToHost(DataTransferMode.EVERY_EXECUTION, deviceOutput); - this.plan = new TornadoExecutionPlan(graph.snapshot()); - } - - private void execute(float[] input, float[] output, int actualBatchSize) { - float[] executionInput = - prepareExecutionInput(input, paddedInput, actualBatchSize, batchSize, cols); - Q4ProjectionKernel.quantize( - executionInput, preparedActivations, preparedScales, batchSize, cols); - deviceActivations.getSegment().copyFrom(MemorySegment.ofArray(preparedActivations)); - deviceScales.getSegment().copyFrom(MemorySegment.ofArray(preparedScales)); - plan.execute(); - copyOutput(deviceOutput, output, actualBatchSize, rows); - } - - @Override - public void close() { - try { - plan.close(); - } catch (Exception exception) { - throw new IllegalStateException("could not close Tornado projection plan", exception); - } - } - } - - private static final class DualProjectionPlan implements AutoCloseable { - private final int batchSize; - private final int firstRows; - private final int secondRows; - private final int cols; - private final float[] paddedInput; - private final byte[] preparedActivations; - private final float[] preparedScales; - private final ByteArray deviceActivations; - private final FloatArray deviceScales; - private final FloatArray firstDeviceOutput; - private final FloatArray secondDeviceOutput; - private final TornadoExecutionPlan plan; - - private DualProjectionPlan( - String name, - MemorySegment firstWeights, - int firstRows, - MemorySegment secondWeights, - int secondRows, - int batchSize, - int cols) { - this.batchSize = batchSize; - this.firstRows = firstRows; - this.secondRows = secondRows; - this.cols = cols; - int activationEntries = Math.multiplyExact(batchSize, cols); - this.paddedInput = new float[activationEntries]; - this.preparedActivations = new byte[activationEntries]; - this.preparedScales = new float[activationEntries / 32]; - ByteArray firstDeviceWeights = ByteArray.fromSegment(firstWeights); - ByteArray secondDeviceWeights = ByteArray.fromSegment(secondWeights); - this.deviceActivations = new ByteArray(activationEntries); - this.deviceScales = new FloatArray(preparedScales.length); - this.firstDeviceOutput = new FloatArray(Math.multiplyExact(batchSize, firstRows)); - this.secondDeviceOutput = new FloatArray(Math.multiplyExact(batchSize, secondRows)); - Q4ProjectionKernel.validate( - firstDeviceWeights, - deviceActivations, - deviceScales, - firstDeviceOutput, - batchSize, - firstRows, - cols); - Q4ProjectionKernel.validate( - secondDeviceWeights, - deviceActivations, - deviceScales, - secondDeviceOutput, - batchSize, - secondRows, - cols); - TaskGraph graph = - new TaskGraph(name) - .transferToDevice( - DataTransferMode.FIRST_EXECUTION, firstDeviceWeights, secondDeviceWeights) - .transferToDevice(DataTransferMode.EVERY_EXECUTION, deviceActivations, deviceScales) - .task( - "multiply-dual", - Q4ProjectionKernel::multiplyDual, - firstDeviceWeights, - firstRows, - secondDeviceWeights, - secondRows, - deviceActivations, - deviceScales, - firstDeviceOutput, - secondDeviceOutput, - batchSize, - cols) - .transferToHost( - DataTransferMode.EVERY_EXECUTION, firstDeviceOutput, secondDeviceOutput); - this.plan = new TornadoExecutionPlan(graph.snapshot()); - } - - private void execute( - float[] input, float[] firstOutput, float[] secondOutput, int actualBatchSize) { - float[] executionInput = - prepareExecutionInput(input, paddedInput, actualBatchSize, batchSize, cols); - prepareAndStage( - executionInput, - preparedActivations, - preparedScales, - deviceActivations, - deviceScales, - batchSize, - cols); - plan.execute(); - copyOutput(firstDeviceOutput, firstOutput, actualBatchSize, firstRows); - copyOutput(secondDeviceOutput, secondOutput, actualBatchSize, secondRows); - } - - @Override - public void close() { - closePlan(plan); - } - } - - private static final class TripleProjectionPlan implements AutoCloseable { - private final int batchSize; - private final int firstRows; - private final int secondRows; - private final int thirdRows; - private final int cols; - private final float[] paddedInput; - private final byte[] preparedActivations; - private final float[] preparedScales; - private final ByteArray deviceActivations; - private final FloatArray deviceScales; - private final FloatArray firstDeviceOutput; - private final FloatArray secondDeviceOutput; - private final FloatArray thirdDeviceOutput; - private final TornadoExecutionPlan plan; - - private TripleProjectionPlan( - String name, - MemorySegment firstWeights, - int firstRows, - MemorySegment secondWeights, - int secondRows, - MemorySegment thirdWeights, - int thirdRows, - int batchSize, - int cols) { - this.batchSize = batchSize; - this.firstRows = firstRows; - this.secondRows = secondRows; - this.thirdRows = thirdRows; - this.cols = cols; - int activationEntries = Math.multiplyExact(batchSize, cols); - this.paddedInput = new float[activationEntries]; - this.preparedActivations = new byte[activationEntries]; - this.preparedScales = new float[activationEntries / 32]; - ByteArray firstDeviceWeights = ByteArray.fromSegment(firstWeights); - ByteArray secondDeviceWeights = ByteArray.fromSegment(secondWeights); - ByteArray thirdDeviceWeights = ByteArray.fromSegment(thirdWeights); - this.deviceActivations = new ByteArray(activationEntries); - this.deviceScales = new FloatArray(preparedScales.length); - this.firstDeviceOutput = new FloatArray(Math.multiplyExact(batchSize, firstRows)); - this.secondDeviceOutput = new FloatArray(Math.multiplyExact(batchSize, secondRows)); - this.thirdDeviceOutput = new FloatArray(Math.multiplyExact(batchSize, thirdRows)); - Q4ProjectionKernel.validate( - firstDeviceWeights, - deviceActivations, - deviceScales, - firstDeviceOutput, - batchSize, - firstRows, - cols); - Q4ProjectionKernel.validate( - secondDeviceWeights, - deviceActivations, - deviceScales, - secondDeviceOutput, - batchSize, - secondRows, - cols); - Q4ProjectionKernel.validate( - thirdDeviceWeights, - deviceActivations, - deviceScales, - thirdDeviceOutput, - batchSize, - thirdRows, - cols); - TaskGraph graph = - new TaskGraph(name) - .transferToDevice( - DataTransferMode.FIRST_EXECUTION, - firstDeviceWeights, - secondDeviceWeights, - thirdDeviceWeights) - .transferToDevice(DataTransferMode.EVERY_EXECUTION, deviceActivations, deviceScales) - .task( - "multiply-triple", - Q4ProjectionKernel::multiplyTriple, - firstDeviceWeights, - firstRows, - secondDeviceWeights, - secondRows, - thirdDeviceWeights, - thirdRows, - deviceActivations, - deviceScales, - firstDeviceOutput, - secondDeviceOutput, - thirdDeviceOutput, - batchSize, - cols) - .transferToHost( - DataTransferMode.EVERY_EXECUTION, - firstDeviceOutput, - secondDeviceOutput, - thirdDeviceOutput); - this.plan = new TornadoExecutionPlan(graph.snapshot()); - } - - private void execute( - float[] input, - float[] firstOutput, - float[] secondOutput, - float[] thirdOutput, - int actualBatchSize) { - float[] executionInput = - prepareExecutionInput(input, paddedInput, actualBatchSize, batchSize, cols); - prepareAndStage( - executionInput, - preparedActivations, - preparedScales, - deviceActivations, - deviceScales, - batchSize, - cols); - plan.execute(); - copyOutput(firstDeviceOutput, firstOutput, actualBatchSize, firstRows); - copyOutput(secondDeviceOutput, secondOutput, actualBatchSize, secondRows); - copyOutput(thirdDeviceOutput, thirdOutput, actualBatchSize, thirdRows); - } - - @Override - public void close() { - closePlan(plan); - } - } - - private record KQuantProjectionKey( - List addresses, - List byteSizes, - List rows, - List types, - int batchSize, - int cols) {} - - /** - * One compiled K-quant dispatch: up to three weight tensors sharing a single Q8_K activation - * preparation. - * - *

The shape is fixed at construction so the task graph and its compiled code are reused across - * calls, exactly like the Q4_0 plans. Only shapes {@link #supportsDual} and {@link - * #supportsTriple} admit can reach here, so the task selection below is total. - */ - private static final class KQuantProjectionPlan implements AutoCloseable { - private final int batchSize; - private final int cols; - private final int[] rows; - private final float[] paddedInput; - private final byte[] preparedActivations; - private final float[] preparedScales; - private final int[] preparedSums; - private final ByteArray deviceActivations; - private final FloatArray deviceScales; - private final IntArray deviceSums; - private final FloatArray[] deviceOutputs; - private final boolean stagesSums; - private final TornadoExecutionPlan plan; - - private KQuantProjectionPlan( - String name, - MemorySegment[] weights, - GgufTensorType[] types, - int[] rows, - int batchSize, - int cols) { - this.batchSize = batchSize; - this.cols = cols; - this.rows = rows.clone(); - int activationEntries = Math.multiplyExact(batchSize, cols); - this.paddedInput = new float[activationEntries]; - this.preparedActivations = new byte[activationEntries]; - this.preparedScales = - new float[activationEntries / KQuantProjectionKernel.SUPER_BLOCK_VALUES]; - this.preparedSums = new int[activationEntries / KQuantProjectionKernel.SUM_BLOCK_VALUES]; - this.deviceActivations = new ByteArray(activationEntries); - this.deviceScales = new FloatArray(preparedScales.length); - this.deviceSums = new IntArray(preparedSums.length); - ByteArray[] deviceWeights = new ByteArray[weights.length]; - this.deviceOutputs = new FloatArray[weights.length]; - for (int index = 0; index < weights.length; index++) { - deviceWeights[index] = ByteArray.fromSegment(weights[index]); - deviceOutputs[index] = new FloatArray(Math.multiplyExact(batchSize, rows[index])); - if (types[index] == GgufTensorType.Q6_K) { - KQuantProjectionKernel.validateQ6K( - deviceWeights[index], - deviceActivations, - deviceScales, - deviceOutputs[index], - batchSize, - rows[index], - cols); - } else { - KQuantProjectionKernel.validateQ4K( - deviceWeights[index], - deviceActivations, - deviceScales, - deviceSums, - deviceOutputs[index], - batchSize, - rows[index], - cols); - } - } - // A Q6_K-only dispatch reads no Q8_K block sums, so the buffer must not be staged for a - // task that never takes it as a parameter. - boolean usesSums = false; - for (GgufTensorType type : types) { - usesSums |= type == GgufTensorType.Q4_K; - } - this.stagesSums = usesSums; - TaskGraph graph = - new TaskGraph(name) - .transferToDevice(DataTransferMode.FIRST_EXECUTION, (Object[]) deviceWeights); - graph = - usesSums - ? graph.transferToDevice( - DataTransferMode.EVERY_EXECUTION, deviceActivations, deviceScales, deviceSums) - : graph.transferToDevice( - DataTransferMode.EVERY_EXECUTION, deviceActivations, deviceScales); - this.plan = - new TornadoExecutionPlan( - addTask(graph, deviceWeights, types, rows) - .transferToHost(DataTransferMode.EVERY_EXECUTION, (Object[]) deviceOutputs) - .snapshot()); - } - - private TaskGraph addTask( - TaskGraph graph, ByteArray[] deviceWeights, GgufTensorType[] types, int[] rows) { - if (deviceWeights.length == 1) { - if (types[0] == GgufTensorType.Q6_K) { - return graph.task( - "multiply-q6k", - KQuantProjectionKernel::multiplyQ6K, - deviceWeights[0], - deviceActivations, - deviceScales, - deviceOutputs[0], - batchSize, - rows[0], - cols); - } - return graph.task( - "multiply-q4k", - KQuantProjectionKernel::multiplyQ4K, - deviceWeights[0], - deviceActivations, - deviceScales, - deviceSums, - deviceOutputs[0], - batchSize, - rows[0], - cols); - } - if (deviceWeights.length == 2) { - return graph.task( - "multiply-q4k-dual", - KQuantProjectionKernel::multiplyQ4KDual, - deviceWeights[0], - rows[0], - deviceWeights[1], - rows[1], - deviceActivations, - deviceScales, - deviceSums, - deviceOutputs[0], - deviceOutputs[1], - batchSize, - cols); - } - if (types[2] == GgufTensorType.Q6_K) { - return graph.task( - "multiply-q4k-q4k-q6k", - KQuantProjectionKernel::multiplyMixedTriple, - deviceWeights[0], - rows[0], - deviceWeights[1], - rows[1], - deviceWeights[2], - rows[2], - deviceActivations, - deviceScales, - deviceSums, - deviceOutputs[0], - deviceOutputs[1], - deviceOutputs[2], - batchSize, - cols); - } - return graph.task( - "multiply-q4k-triple", - KQuantProjectionKernel::multiplyQ4KTriple, - deviceWeights[0], - rows[0], - deviceWeights[1], - rows[1], - deviceWeights[2], - rows[2], - deviceActivations, - deviceScales, - deviceSums, - deviceOutputs[0], - deviceOutputs[1], - deviceOutputs[2], - batchSize, - cols); - } - - private void execute(float[] input, float[][] outputs, int actualBatchSize) { - float[] executionInput = - prepareExecutionInput(input, paddedInput, actualBatchSize, batchSize, cols); - KQuantProjectionKernel.quantize( - executionInput, preparedActivations, preparedScales, preparedSums, batchSize, cols); - deviceActivations.getSegment().copyFrom(MemorySegment.ofArray(preparedActivations)); - deviceScales.getSegment().copyFrom(MemorySegment.ofArray(preparedScales)); - if (stagesSums) { - deviceSums.getSegment().copyFrom(MemorySegment.ofArray(preparedSums)); - } - plan.execute(); - for (int index = 0; index < outputs.length; index++) { - copyOutput(deviceOutputs[index], outputs[index], actualBatchSize, rows[index]); - } - } - - @Override - public void close() { - closePlan(plan); - } - } - - private static void prepareAndStage( - float[] input, - byte[] preparedActivations, - float[] preparedScales, - ByteArray deviceActivations, - FloatArray deviceScales, - int batchSize, - int cols) { - Q4ProjectionKernel.quantize(input, preparedActivations, preparedScales, batchSize, cols); - deviceActivations.getSegment().copyFrom(MemorySegment.ofArray(preparedActivations)); - deviceScales.getSegment().copyFrom(MemorySegment.ofArray(preparedScales)); - } - - static float[] prepareExecutionInput( - float[] input, float[] paddedInput, int actualBatchSize, int executionBatchSize, int cols) { - int actualEntries = Math.multiplyExact(actualBatchSize, cols); - int executionEntries = Math.multiplyExact(executionBatchSize, cols); - if (input.length < actualEntries || paddedInput.length != executionEntries) { - throw new IllegalArgumentException("input storage does not match execution batch shape"); - } - System.arraycopy(input, 0, paddedInput, 0, actualEntries); - Arrays.fill(paddedInput, actualEntries, executionEntries, 0.0f); - return paddedInput; - } - - private static void copyOutput(FloatArray deviceOutput, float[] output, int batchSize, int rows) { - long byteSize = Math.multiplyExact((long) batchSize * rows, Float.BYTES); - MemorySegment.ofArray(output) - .asSlice(0, byteSize) - .copyFrom(deviceOutput.getSegment().asSlice(0, byteSize)); - } - - private static void closePlan(TornadoExecutionPlan plan) { - try { - plan.close(); - } catch (Exception exception) { - throw new IllegalStateException("could not close Tornado projection plan", exception); - } - } -} diff --git a/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/TornadoPureJavaBackendProvider.java b/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/TornadoPureJavaBackendProvider.java deleted file mode 100644 index 4937bedb..00000000 --- a/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/TornadoPureJavaBackendProvider.java +++ /dev/null @@ -1,72 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.backend.tornado; - -import com.integrallis.models.api.BackendConfiguration; -import com.integrallis.models.backend.purejava.PureJavaBackend; -import com.integrallis.models.backend.purejava.spi.PureJavaBackendProvider; -import java.io.IOException; -import java.nio.file.Files; -import java.nio.file.Path; -import java.util.List; -import java.util.Optional; - -/** Service-loaded Tornado provider used by Models and ModelJars automatic backend loading. */ -public final class TornadoPureJavaBackendProvider implements PureJavaBackendProvider { - - @Override - public Optional tryLoad( - Path modelPath, BackendConfiguration backendConfiguration) { - Path model = modelPath.toAbsolutePath().normalize(); - if (!Files.isRegularFile(model)) { - return Optional.empty(); - } - List devices; - try { - devices = TornadoRuntimeDevices.discover(); - } catch (LinkageError | RuntimeException failure) { - return Optional.empty(); - } - long modelBytes; - try { - modelBytes = Files.size(model); - } catch (IOException failure) { - return Optional.empty(); - } - TornadoBackendOptions options = TornadoBackendOptions.defaults(); - AcceleratorEligibility.Decision decision = - AcceleratorEligibility.select( - devices, - DeviceMemoryRequest.ofModelFile( - modelLabel(model), modelBytes, options.accelerateDecode())); - if (!decision.eligible()) { - return Optional.empty(); - } - TornadoBackendOptions required = - new TornadoBackendOptions( - options.accelerateDecode(), - options.eagerReadiness(), - true, - options.executionBatchSize()); - TornadoBackendRuntime runtime = TornadoBackend.open(model, backendConfiguration, required); - return Optional.of(runtime.detachBackend()); - } - - private static String modelLabel(Path model) { - Path fileName = model.getFileName(); - return fileName == null ? model.toString() : fileName.toString(); - } -} diff --git a/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/TornadoRuntimeDevices.java b/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/TornadoRuntimeDevices.java deleted file mode 100644 index 90200047..00000000 --- a/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/TornadoRuntimeDevices.java +++ /dev/null @@ -1,52 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.backend.tornado; - -import java.util.ArrayList; -import java.util.List; -import uk.ac.manchester.tornado.api.TornadoBackend; -import uk.ac.manchester.tornado.api.TornadoRuntime; -import uk.ac.manchester.tornado.api.common.TornadoDevice; -import uk.ac.manchester.tornado.api.runtime.TornadoRuntimeProvider; - -/** Discovers accelerator devices from the active TornadoVM runtime. */ -public final class TornadoRuntimeDevices { - private TornadoRuntimeDevices() {} - - public static List discover() { - try { - TornadoRuntime runtime = TornadoRuntimeProvider.getTornadoRuntime(); - List devices = new ArrayList<>(); - for (int backendIndex = 0; backendIndex < runtime.getNumBackends(); backendIndex++) { - TornadoBackend backend = runtime.getBackend(backendIndex); - String backendType = runtime.getBackendType(backendIndex).name(); - for (int deviceIndex = 0; deviceIndex < backend.getNumDevices(); deviceIndex++) { - TornadoDevice device = backend.getDevice(deviceIndex); - devices.add( - new AcceleratorEligibility.DeviceCapabilities( - device.getDeviceName(), - backendType, - device.getDeviceType().name(), - device.getMaxGlobalMemory(), - device.getMaxAllocMemory())); - } - } - return List.copyOf(devices); - } catch (LinkageError | RuntimeException unavailable) { - return List.of(); - } - } -} diff --git a/backend-tornado/src/main/resources/META-INF/services/com.integrallis.models.backend.purejava.spi.PureJavaBackendProvider b/backend-tornado/src/main/resources/META-INF/services/com.integrallis.models.backend.purejava.spi.PureJavaBackendProvider deleted file mode 100644 index adddedcd..00000000 --- a/backend-tornado/src/main/resources/META-INF/services/com.integrallis.models.backend.purejava.spi.PureJavaBackendProvider +++ /dev/null @@ -1 +0,0 @@ -com.integrallis.models.backend.tornado.TornadoPureJavaBackendProvider diff --git a/backend-tornado/src/test/java/com/integrallis/models/backend/tornado/AcceleratorEligibilityTest.java b/backend-tornado/src/test/java/com/integrallis/models/backend/tornado/AcceleratorEligibilityTest.java deleted file mode 100644 index 7ad81d3e..00000000 --- a/backend-tornado/src/test/java/com/integrallis/models/backend/tornado/AcceleratorEligibilityTest.java +++ /dev/null @@ -1,80 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.backend.tornado; - -import static org.assertj.core.api.Assertions.assertThat; - -import java.util.List; -import org.junit.jupiter.api.Test; - -class AcceleratorEligibilityTest { - private static final long MIB = 1024L * 1024L; - private static final long GIB = 1024L * MIB; - - @Test - void selectsAQualifiedPtxGpuWithEnoughMemoryForPrefillAndDecodePlans() { - AcceleratorEligibility.Decision decision = - AcceleratorEligibility.select( - List.of( - device("host CPU", "JAVA", "CPU", 16 * GIB, 4 * GIB), - device("NVIDIA A16", "PTX", "GPU", 2 * GIB, 512 * MIB)), - 410 * MIB, - true); - - assertThat(decision.eligible()).isTrue(); - assertThat(decision.device().name()).isEqualTo("NVIDIA A16"); - assertThat(decision.requiredBytes()).isEqualTo(1_076 * MIB); - } - - @Test - void fallsBackWhenTheQualifiedDeviceCannotHoldTheRetainedPlansWithHeadroom() { - AcceleratorEligibility.Decision decision = - AcceleratorEligibility.select( - List.of(device("NVIDIA A16", "PTX", "GPU", 2 * GIB, 512 * MIB)), 900 * MIB, true); - - assertThat(decision.eligible()).isFalse(); - assertThat(decision.reason()).contains("NVIDIA A16 has 1.50 GiB usable of 2.00 GiB"); - assertThat(decision.reason()).contains("needs 2.01 GiB"); - assertThat(decision.reason()).contains("weights 1.76 GiB (PER_SHAPE_WHOLE_MODEL)"); - } - - @Test - void leavesUnqualifiedVendorsAndCpuDevicesOnTheJavaFallback() { - AcceleratorEligibility.Decision decision = - AcceleratorEligibility.select( - List.of( - device("AMD Radeon", "OPENCL", "GPU", 8 * GIB, 2 * GIB), - device("Intel Xeon", "OPENCL", "CPU", 64 * GIB, 16 * GIB)), - 410 * MIB, - false); - - assertThat(decision.eligible()).isFalse(); - assertThat(decision.reason()).contains("PTX or CUDA backend on a GPU device"); - } - - @Test - void rejectsInvalidModelSizesWithoutInspectingDevices() { - AcceleratorEligibility.Decision decision = AcceleratorEligibility.select(List.of(), 0, false); - - assertThat(decision.eligible()).isFalse(); - assertThat(decision.reason()).contains("model size"); - } - - private static AcceleratorEligibility.DeviceCapabilities device( - String name, String backend, String type, long memory, long allocation) { - return new AcceleratorEligibility.DeviceCapabilities(name, backend, type, memory, allocation); - } -} diff --git a/backend-tornado/src/test/java/com/integrallis/models/backend/tornado/KQuantProjectionKernelTest.java b/backend-tornado/src/test/java/com/integrallis/models/backend/tornado/KQuantProjectionKernelTest.java deleted file mode 100644 index 1af78ed2..00000000 --- a/backend-tornado/src/test/java/com/integrallis/models/backend/tornado/KQuantProjectionKernelTest.java +++ /dev/null @@ -1,697 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.backend.tornado; - -import static org.assertj.core.api.Assertions.assertThat; -import static org.assertj.core.api.Assertions.assertThatThrownBy; -import static org.assertj.core.api.Assertions.within; - -import com.integrallis.vectors.core.VectorUtil; -import java.lang.foreign.MemorySegment; -import java.util.Random; -import org.junit.jupiter.api.Test; -import uk.ac.manchester.tornado.api.types.arrays.ByteArray; -import uk.ac.manchester.tornado.api.types.arrays.FloatArray; -import uk.ac.manchester.tornado.api.types.arrays.IntArray; - -/** - * Off-device parity tests for the K-quant TornadoVM kernels. - * - *

These run the kernels as ordinary sequential Java — {@code @Parallel} is an annotation, so a - * direct call from a unit test executes the same arithmetic the PTX backend compiles — and score - * them against the production vectors-core CPU kernels the pure-Java backend actually uses. - * - *

Two numeric contracts are asserted, deliberately separately: - * - *

    - *
  1. Exact. On super-blocks whose scales are powers of two and whose quantized sums stay - * well inside the exactly representable float range, the kernel must equal the CPU kernel - * bit-for-bit. This is the strong statement: it proves the GGUF super-block decode — the - * six-bit scale/min packing, the nibble group order, the Q6_K two-bit high pairs, and the - * Q8_K per-16 sums — is identical, because every integer reduction is exact and any decode - * error would move the result by at least one quantization step. - *
  2. Bounded. On pseudo-random super-blocks the kernel cannot be exact: the CPU kernels - * fuse the per-block scale application with {@code Math.fma} and this kernel uses a plain - * multiply and add, which is the operation set the existing Q4_0 kernel is known to compile - * to PTX with. Q4_K applies two such steps per super-block and Q6_K one, so the guaranteed - * contract is at most two float roundings per super-block: {@code |kernel - cpu| <= 2 - * * (cols / 256) * ulp(max |cpu|)}. That is what {@link #assertWithinFusedMultiplyAddBudget} - * asserts. - *

    Measured here on 2026-09-18 (Apple M-series, vectors-core 0.1.22 Panama provider, - * 256-bit species, {@code fastVectorFMA=true}), over 40 pseudo-random matrices per width at - * batch 4 by 8 rows: Q4_K worst absolute difference 7.6e-6 at cols=256, 6.1e-5 at cols=1024, - * 1.2e-4 at cols=4096; Q6_K 0.0, 9.2e-5, 1.8e-4 for the same widths, against reference - * magnitudes of 3.3e2, 5.3e2 and 1.4e3. Every one of those is inside the asserted budget with - * at least an order of magnitude to spare, and all of them scale with the super-block count - * exactly as a per-super-block rounding predicts. - *

- * - *

What these tests cannot establish: that TornadoVM's PTX backend lowers this bytecode to - * the same arithmetic. A device run can differ by contracting {@code a * b + c} into a hardware - * fused multiply-add, which would move results toward the CPU kernel rather than away from it, but - * that is an argument, not a measurement. The device-side statement needs a GPU host. - */ -class KQuantProjectionKernelTest { - private static final int SUPER_BLOCK = KQuantProjectionKernel.SUPER_BLOCK_VALUES; - private static final int SUM_BLOCK = KQuantProjectionKernel.SUM_BLOCK_VALUES; - - // --- Q4_K --- - - @Test - void matchesProductionQ4KProjectionAcrossBatches() { - int batchSize = 3; - int rows = 9; - int cols = 512; - byte[] weights = randomQ4KMatrix(rows, cols, 31L); - float[] input = randomFloats(batchSize * cols, 37L); - - float[] actual = runQ4K(weights, input, batchSize, rows, cols); - - assertWithinFusedMultiplyAddBudget( - actual, referenceQ4K(weights, input, batchSize, rows, cols), cols); - } - - @Test - void matchesProductionQ4KProjectionExactlyOnExactlyRepresentableBlocks() { - int batchSize = 2; - int rows = 5; - int cols = 512; - byte[] weights = exactQ4KMatrix(rows, cols, 71L); - float[] input = exactFloats(batchSize, cols, 73L); - - float[] actual = runQ4K(weights, input, batchSize, rows, cols); - - assertThat(actual).isEqualTo(referenceQ4K(weights, input, batchSize, rows, cols)); - } - - @Test - void matchesTwoProductionQ4KProjectionsWithOnePreparedActivation() { - int batchSize = 2; - int firstRows = 7; - int secondRows = 5; - int cols = 256; - byte[] firstWeights = exactQ4KMatrix(firstRows, cols, 41L); - byte[] secondWeights = exactQ4KMatrix(secondRows, cols, 43L); - float[] input = exactFloats(batchSize, cols, 47L); - Q8KActivation activation = Q8KActivation.of(input, batchSize, cols); - FloatArray firstOutput = new FloatArray(batchSize * firstRows); - FloatArray secondOutput = new FloatArray(batchSize * secondRows); - - KQuantProjectionKernel.multiplyQ4KDual( - ByteArray.fromArray(firstWeights), - firstRows, - ByteArray.fromArray(secondWeights), - secondRows, - activation.quants(), - activation.scales(), - activation.sums(), - firstOutput, - secondOutput, - batchSize, - cols); - - assertThat(firstOutput.toHeapArray()) - .isEqualTo(referenceQ4K(firstWeights, input, batchSize, firstRows, cols)); - assertThat(secondOutput.toHeapArray()) - .isEqualTo(referenceQ4K(secondWeights, input, batchSize, secondRows, cols)); - } - - @Test - void matchesThreeProductionQ4KProjectionsWithOnePreparedActivation() { - int batchSize = 2; - int firstRows = 4; - int secondRows = 3; - int thirdRows = 2; - int cols = 256; - byte[] firstWeights = randomQ4KMatrix(firstRows, cols, 53L); - byte[] secondWeights = randomQ4KMatrix(secondRows, cols, 59L); - byte[] thirdWeights = randomQ4KMatrix(thirdRows, cols, 61L); - float[] input = randomFloats(batchSize * cols, 67L); - Q8KActivation activation = Q8KActivation.of(input, batchSize, cols); - FloatArray firstOutput = new FloatArray(batchSize * firstRows); - FloatArray secondOutput = new FloatArray(batchSize * secondRows); - FloatArray thirdOutput = new FloatArray(batchSize * thirdRows); - - KQuantProjectionKernel.multiplyQ4KTriple( - ByteArray.fromArray(firstWeights), - firstRows, - ByteArray.fromArray(secondWeights), - secondRows, - ByteArray.fromArray(thirdWeights), - thirdRows, - activation.quants(), - activation.scales(), - activation.sums(), - firstOutput, - secondOutput, - thirdOutput, - batchSize, - cols); - - assertWithinFusedMultiplyAddBudget( - firstOutput.toHeapArray(), - referenceQ4K(firstWeights, input, batchSize, firstRows, cols), - cols); - assertWithinFusedMultiplyAddBudget( - secondOutput.toHeapArray(), - referenceQ4K(secondWeights, input, batchSize, secondRows, cols), - cols); - assertWithinFusedMultiplyAddBudget( - thirdOutput.toHeapArray(), - referenceQ4K(thirdWeights, input, batchSize, thirdRows, cols), - cols); - } - - // --- Q6_K --- - - @Test - void matchesProductionQ6KProjectionAcrossBatches() { - int batchSize = 3; - int rows = 9; - int cols = 512; - byte[] weights = randomQ6KMatrix(rows, cols, 83L); - float[] input = randomFloats(batchSize * cols, 89L); - - float[] actual = runQ6K(weights, input, batchSize, rows, cols); - - assertWithinFusedMultiplyAddBudget( - actual, referenceQ6K(weights, input, batchSize, rows, cols), cols); - } - - @Test - void matchesProductionQ6KProjectionExactlyOnExactlyRepresentableBlocks() { - int batchSize = 2; - int rows = 5; - int cols = 512; - byte[] weights = exactQ6KMatrix(rows, cols, 97L); - float[] input = exactFloats(batchSize, cols, 101L); - - float[] actual = runQ6K(weights, input, batchSize, rows, cols); - - assertThat(actual).isEqualTo(referenceQ6K(weights, input, batchSize, rows, cols)); - } - - // --- Mixed Q4_K / Q4_K / Q6_K, the Q4_K_M attention projection group --- - - @Test - void matchesTheProductionMixedAttentionProjectionGroup() { - int batchSize = 2; - int queryRows = 6; - int keyRows = 3; - int valueRows = 3; - int cols = 256; - byte[] queryWeights = exactQ4KMatrix(queryRows, cols, 103L); - byte[] keyWeights = exactQ4KMatrix(keyRows, cols, 107L); - byte[] valueWeights = exactQ6KMatrix(valueRows, cols, 109L); - float[] input = exactFloats(batchSize, cols, 113L); - Q8KActivation activation = Q8KActivation.of(input, batchSize, cols); - FloatArray queryOutput = new FloatArray(batchSize * queryRows); - FloatArray keyOutput = new FloatArray(batchSize * keyRows); - FloatArray valueOutput = new FloatArray(batchSize * valueRows); - - KQuantProjectionKernel.multiplyMixedTriple( - ByteArray.fromArray(queryWeights), - queryRows, - ByteArray.fromArray(keyWeights), - keyRows, - ByteArray.fromArray(valueWeights), - valueRows, - activation.quants(), - activation.scales(), - activation.sums(), - queryOutput, - keyOutput, - valueOutput, - batchSize, - cols); - - assertThat(queryOutput.toHeapArray()) - .isEqualTo(referenceQ4K(queryWeights, input, batchSize, queryRows, cols)); - assertThat(keyOutput.toHeapArray()) - .isEqualTo(referenceQ4K(keyWeights, input, batchSize, keyRows, cols)); - assertThat(valueOutput.toHeapArray()) - .isEqualTo(referenceQ6K(valueWeights, input, batchSize, valueRows, cols)); - } - - // --- Q8_K activation preparation --- - - @Test - void preparesQ8KActivationsWithBlockSumsThatMatchTheStoredQuants() { - int batchSize = 2; - int cols = 512; - float[] input = randomFloats(batchSize * cols, 127L); - byte[] quants = new byte[batchSize * cols]; - float[] scales = new float[batchSize * cols / SUPER_BLOCK]; - int[] sums = new int[batchSize * cols / SUM_BLOCK]; - - KQuantProjectionKernel.quantize(input, quants, scales, sums, batchSize, cols); - - for (int sumBlock = 0; sumBlock < sums.length; sumBlock++) { - int expected = 0; - for (int index = 0; index < SUM_BLOCK; index++) { - expected += quants[sumBlock * SUM_BLOCK + index]; - } - assertThat(sums[sumBlock]).as("sum block %d", sumBlock).isEqualTo(expected); - assertThat(expected).isBetween(-Short.MAX_VALUE, (int) Short.MAX_VALUE); - } - // Q8_K scales carry the sign of the block extremum, unlike Q8_0's absolute-maximum scale. - for (float scale : scales) { - assertThat(scale).isNotZero().isFinite(); - } - for (int index = 0; index < quants.length; index++) { - assertThat((int) quants[index]).isBetween(-127, 127); - float reconstructed = quants[index] * scales[index / SUPER_BLOCK]; - assertThat(reconstructed) - .as("reconstructed element %d", index) - .isCloseTo(input[index], within(0.02f)); - } - } - - @Test - void preparesZeroActivationBlocksWithoutDividingByZero() { - int cols = 256; - byte[] quants = new byte[cols]; - float[] scales = new float[1]; - int[] sums = new int[cols / SUM_BLOCK]; - - KQuantProjectionKernel.quantize(new float[cols], quants, scales, sums, 1, cols); - - assertThat(scales[0]).isZero(); - assertThat(quants).containsOnly((byte) 0); - assertThat(sums).containsOnly(0); - } - - @Test - void projectsZeroActivationsToZero() { - int rows = 8; - int cols = 256; - byte[] weights = randomQ4KMatrix(rows, cols, 131L); - - float[] actual = runQ4K(weights, new float[cols], 1, rows, cols); - - assertThat(actual).containsOnly(0.0f); - } - - // --- Validation --- - - @Test - void rejectsQ4KActivationSumStorageThatDoesNotMatchTheShape() { - assertThatThrownBy( - () -> - KQuantProjectionKernel.validateQ4K( - new ByteArray(KQuantProjectionKernel.Q4_K_BLOCK_BYTES), - new ByteArray(SUPER_BLOCK), - new FloatArray(1), - new IntArray(3), - new FloatArray(1), - 1, - 1, - SUPER_BLOCK)) - .isInstanceOf(IllegalArgumentException.class) - .hasMessageContaining("activation sums"); - } - - @Test - void rejectsQ4KShapesThatAreNotWholeSuperBlocks() { - assertThatThrownBy( - () -> - KQuantProjectionKernel.validateQ4K( - new ByteArray(KQuantProjectionKernel.Q4_K_BLOCK_BYTES), - new ByteArray(128), - new FloatArray(1), - new IntArray(8), - new FloatArray(1), - 1, - 1, - 128)) - .isInstanceOf(IllegalArgumentException.class) - .hasMessageContaining("multiple of 256"); - } - - @Test - void rejectsQ4KWeightsThatDoNotMatchTheShape() { - assertThatThrownBy( - () -> - KQuantProjectionKernel.validateQ4K( - new ByteArray(KQuantProjectionKernel.Q4_K_BLOCK_BYTES + 1), - new ByteArray(SUPER_BLOCK), - new FloatArray(1), - new IntArray(16), - new FloatArray(1), - 1, - 1, - SUPER_BLOCK)) - .isInstanceOf(IllegalArgumentException.class) - .hasMessageContaining("weights"); - } - - @Test - void rejectsQ4KActivationStorageThatDoesNotMatchTheShape() { - assertThatThrownBy( - () -> - KQuantProjectionKernel.validateQ4K( - new ByteArray(KQuantProjectionKernel.Q4_K_BLOCK_BYTES), - new ByteArray(SUPER_BLOCK - 1), - new FloatArray(1), - new IntArray(16), - new FloatArray(1), - 1, - 1, - SUPER_BLOCK)) - .isInstanceOf(IllegalArgumentException.class) - .hasMessageContaining("activations"); - } - - @Test - void rejectsQ4KActivationScaleStorageThatDoesNotMatchTheShape() { - assertThatThrownBy( - () -> - KQuantProjectionKernel.validateQ4K( - new ByteArray(KQuantProjectionKernel.Q4_K_BLOCK_BYTES), - new ByteArray(SUPER_BLOCK), - new FloatArray(2), - new IntArray(16), - new FloatArray(1), - 1, - 1, - SUPER_BLOCK)) - .isInstanceOf(IllegalArgumentException.class) - .hasMessageContaining("activation scales"); - } - - @Test - void rejectsQ6KOutputStorageThatDoesNotMatchTheShape() { - assertThatThrownBy( - () -> - KQuantProjectionKernel.validateQ6K( - new ByteArray(KQuantProjectionKernel.Q6_K_BLOCK_BYTES), - new ByteArray(SUPER_BLOCK), - new FloatArray(1), - new FloatArray(2), - 1, - 1, - SUPER_BLOCK)) - .isInstanceOf(IllegalArgumentException.class) - .hasMessageContaining("output"); - } - - @Test - void rejectsNonPositiveQ6KShapes() { - assertThatThrownBy( - () -> - KQuantProjectionKernel.validateQ6K( - new ByteArray(KQuantProjectionKernel.Q6_K_BLOCK_BYTES), - new ByteArray(SUPER_BLOCK), - new FloatArray(1), - new FloatArray(1), - 0, - 1, - SUPER_BLOCK)) - .isInstanceOf(IllegalArgumentException.class) - .hasMessageContaining("batchSize"); - assertThatThrownBy( - () -> - KQuantProjectionKernel.validateQ6K( - new ByteArray(KQuantProjectionKernel.Q6_K_BLOCK_BYTES), - new ByteArray(SUPER_BLOCK), - new FloatArray(1), - new FloatArray(1), - 1, - 0, - SUPER_BLOCK)) - .isInstanceOf(IllegalArgumentException.class) - .hasMessageContaining("rows"); - } - - @Test - void acceptsWellFormedQ6KStorage() { - KQuantProjectionKernel.validateQ6K( - new ByteArray(KQuantProjectionKernel.Q6_K_BLOCK_BYTES), - new ByteArray(SUPER_BLOCK), - new FloatArray(1), - new FloatArray(1), - 1, - 1, - SUPER_BLOCK); - } - - // --- Harness --- - - private record Q8KActivation(ByteArray quants, FloatArray scales, IntArray sums) { - static Q8KActivation of(float[] input, int batchSize, int cols) { - byte[] quants = new byte[batchSize * cols]; - float[] scales = new float[batchSize * cols / SUPER_BLOCK]; - int[] sums = new int[batchSize * cols / SUM_BLOCK]; - KQuantProjectionKernel.quantize(input, quants, scales, sums, batchSize, cols); - return new Q8KActivation( - ByteArray.fromArray(quants), FloatArray.fromArray(scales), IntArray.fromArray(sums)); - } - } - - private static float[] runQ4K(byte[] weights, float[] input, int batchSize, int rows, int cols) { - Q8KActivation activation = Q8KActivation.of(input, batchSize, cols); - FloatArray output = new FloatArray(batchSize * rows); - ByteArray deviceWeights = ByteArray.fromArray(weights); - KQuantProjectionKernel.validateQ4K( - deviceWeights, - activation.quants(), - activation.scales(), - activation.sums(), - output, - batchSize, - rows, - cols); - KQuantProjectionKernel.multiplyQ4K( - deviceWeights, - activation.quants(), - activation.scales(), - activation.sums(), - output, - batchSize, - rows, - cols); - return output.toHeapArray(); - } - - private static float[] runQ6K(byte[] weights, float[] input, int batchSize, int rows, int cols) { - Q8KActivation activation = Q8KActivation.of(input, batchSize, cols); - FloatArray output = new FloatArray(batchSize * rows); - ByteArray deviceWeights = ByteArray.fromArray(weights); - KQuantProjectionKernel.validateQ6K( - deviceWeights, activation.quants(), activation.scales(), output, batchSize, rows, cols); - KQuantProjectionKernel.multiplyQ6K( - deviceWeights, activation.quants(), activation.scales(), output, batchSize, rows, cols); - return output.toHeapArray(); - } - - private static float[] referenceQ4K( - byte[] weights, float[] input, int batchSize, int rows, int cols) { - float[] output = new float[batchSize * rows]; - VectorUtil.ggufQ4_KQ8_KBatchedMatmul( - input, - MemorySegment.ofArray(weights), - batchSize, - rows, - cols, - output, - new byte[batchSize * cols], - new float[batchSize * cols / SUPER_BLOCK], - new short[batchSize * cols / SUM_BLOCK]); - return output; - } - - private static float[] referenceQ6K( - byte[] weights, float[] input, int batchSize, int rows, int cols) { - float[] output = new float[batchSize * rows]; - VectorUtil.ggufQ6_KQ8_KBatchedMatmul( - input, - MemorySegment.ofArray(weights), - batchSize, - rows, - cols, - output, - new byte[batchSize * cols], - new float[batchSize * cols / SUPER_BLOCK]); - return output; - } - - /** - * Asserts the guaranteed numeric contract: the kernel differs from the production CPU kernel by - * at most two float roundings per super-block, scaled to the magnitude of the reference result. - */ - private static void assertWithinFusedMultiplyAddBudget( - float[] actual, float[] expected, int cols) { - assertThat(actual).hasSameSizeAs(expected); - float magnitude = 0.0f; - for (float value : expected) { - magnitude = Math.max(magnitude, Math.abs(value)); - } - float tolerance = 2.0f * (cols / SUPER_BLOCK) * Math.ulp(magnitude); - for (int index = 0; index < expected.length; index++) { - assertThat(actual[index]) - .as("element %d within %s of %s", index, tolerance, expected[index]) - .isCloseTo(expected[index], within(tolerance)); - } - } - - /** Pseudo-random Q4_K super-blocks spanning the whole six-bit scale and minimum range. */ - private static byte[] randomQ4KMatrix(int rows, int cols, long seed) { - Random random = new Random(seed); - int blocks = rows * cols / SUPER_BLOCK; - byte[] weights = new byte[blocks * KQuantProjectionKernel.Q4_K_BLOCK_BYTES]; - int[] scales = new int[8]; - int[] minimums = new int[8]; - for (int block = 0; block < blocks; block++) { - for (int group = 0; group < 8; group++) { - scales[group] = random.nextInt(64); - minimums[group] = random.nextInt(64); - } - writeQ4KBlock( - weights, - block, - Float.floatToFloat16(0.002f + random.nextFloat() * 0.02f), - Float.floatToFloat16(0.001f + random.nextFloat() * 0.01f), - scales, - minimums, - random); - } - return weights; - } - - /** - * Q4_K super-blocks whose contribution is exactly representable: the two scales are powers of two - * and the six-bit group scales and minimums stay small enough that no partial sum leaves the - * exactly representable integer range of a float. - */ - private static byte[] exactQ4KMatrix(int rows, int cols, long seed) { - Random random = new Random(seed); - int blocks = rows * cols / SUPER_BLOCK; - byte[] weights = new byte[blocks * KQuantProjectionKernel.Q4_K_BLOCK_BYTES]; - int[] scales = new int[8]; - int[] minimums = new int[8]; - for (int block = 0; block < blocks; block++) { - for (int group = 0; group < 8; group++) { - scales[group] = 1 + random.nextInt(3); - minimums[group] = random.nextInt(4); - } - writeQ4KBlock( - weights, - block, - Float.floatToFloat16(1.0f), - Float.floatToFloat16(0.25f), - scales, - minimums, - random); - } - return weights; - } - - private static void writeQ4KBlock( - byte[] weights, - int block, - short delta, - short deltaMin, - int[] scales, - int[] minimums, - Random random) { - int offset = block * KQuantProjectionKernel.Q4_K_BLOCK_BYTES; - weights[offset] = (byte) delta; - weights[offset + 1] = (byte) (delta >>> 8); - weights[offset + 2] = (byte) deltaMin; - weights[offset + 3] = (byte) (deltaMin >>> 8); - for (int group = 0; group < 4; group++) { - weights[offset + 4 + group] = - (byte) ((scales[group] & 0x3F) | ((scales[group + 4] >>> 4) << 6)); - weights[offset + 8 + group] = - (byte) ((minimums[group] & 0x3F) | ((minimums[group + 4] >>> 4) << 6)); - weights[offset + 12 + group] = - (byte) ((scales[group + 4] & 0x0F) | ((minimums[group + 4] & 0x0F) << 4)); - } - for (int index = 0; index < 128; index++) { - weights[offset + 16 + index] = (byte) random.nextInt(256); - } - } - - /** Pseudo-random Q6_K super-blocks spanning the whole signed int8 group-scale range. */ - private static byte[] randomQ6KMatrix(int rows, int cols, long seed) { - Random random = new Random(seed); - int blocks = rows * cols / SUPER_BLOCK; - byte[] weights = new byte[blocks * KQuantProjectionKernel.Q6_K_BLOCK_BYTES]; - for (int block = 0; block < blocks; block++) { - int offset = block * KQuantProjectionKernel.Q6_K_BLOCK_BYTES; - for (int index = 0; index < 192; index++) { - weights[offset + index] = (byte) random.nextInt(256); - } - for (int index = 0; index < 16; index++) { - weights[offset + 192 + index] = (byte) (random.nextInt(65) - 32); - } - short delta = Float.floatToFloat16(0.002f + random.nextFloat() * 0.02f); - weights[offset + 208] = (byte) delta; - weights[offset + 209] = (byte) (delta >>> 8); - } - return weights; - } - - /** Q6_K super-blocks with a power-of-two delta and small group scales. */ - private static byte[] exactQ6KMatrix(int rows, int cols, long seed) { - Random random = new Random(seed); - int blocks = rows * cols / SUPER_BLOCK; - byte[] weights = new byte[blocks * KQuantProjectionKernel.Q6_K_BLOCK_BYTES]; - for (int block = 0; block < blocks; block++) { - int offset = block * KQuantProjectionKernel.Q6_K_BLOCK_BYTES; - for (int index = 0; index < 192; index++) { - weights[offset + index] = (byte) random.nextInt(256); - } - for (int index = 0; index < 16; index++) { - weights[offset + 192 + index] = (byte) (random.nextInt(5) - 2); - } - short delta = Float.floatToFloat16(0.5f); - weights[offset + 208] = (byte) delta; - weights[offset + 209] = (byte) (delta >>> 8); - } - return weights; - } - - private static float[] randomFloats(int length, long seed) { - Random random = new Random(seed); - float[] values = new float[length]; - for (int index = 0; index < length; index++) { - values[index] = random.nextFloat(-2.0f, 2.0f); - } - return values; - } - - /** - * Activations that quantize to Q8_K without rounding error: every value is a small integer times - * 2^-6, and each super-block's extremum is exactly -127 * 2^-6, so the Q8_K scale is 2^-6 and - * each stored quant is the integer it started from. - */ - private static float[] exactFloats(int batchSize, int cols, long seed) { - Random random = new Random(seed); - float[] values = new float[batchSize * cols]; - float step = 1.0f / 64.0f; - for (int batch = 0; batch < batchSize; batch++) { - for (int block = 0; block < cols / SUPER_BLOCK; block++) { - int offset = batch * cols + block * SUPER_BLOCK; - values[offset] = -127.0f * step; - for (int index = 1; index < SUPER_BLOCK; index++) { - values[offset + index] = (random.nextInt(7) - 3) * step; - } - } - } - return values; - } -} diff --git a/backend-tornado/src/test/java/com/integrallis/models/backend/tornado/LargeModelEligibilityTest.java b/backend-tornado/src/test/java/com/integrallis/models/backend/tornado/LargeModelEligibilityTest.java deleted file mode 100644 index be15d775..00000000 --- a/backend-tornado/src/test/java/com/integrallis/models/backend/tornado/LargeModelEligibilityTest.java +++ /dev/null @@ -1,292 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.backend.tornado; - -import static org.assertj.core.api.Assertions.assertThat; - -import java.util.List; -import java.util.stream.Stream; -import org.junit.jupiter.api.Test; -import org.junit.jupiter.params.ParameterizedTest; -import org.junit.jupiter.params.provider.Arguments; -import org.junit.jupiter.params.provider.MethodSource; - -/** - * Device-independent capacity arithmetic for 27B-class Q4_K_M models. - * - *

Every model constant below is read from this repository, not assumed: - * - *

    - *
  • weights {@code 1_650_027_640 + 15_130_165_248} are the resident and routed-expert byte - * counts asserted in {@code Gemma4LargeModelFixtureSlowTest} for the pinned {@code - * gemma-4-26B-A4B-it-Q4_K_M.gguf}; - *
  • the KV geometry (30 layers, 25 sliding with keyDim 2048, 5 full with keyDim 1024, ring - * capacity {@code slidingWindow 1024 + prefillCapacity 1}) is asserted in {@code - * Gemma4ConfigTest.parsesThePinnedGemma426BA4BMetadata} and built in {@code - * Gemma4KvCache.create}; - *
  • the retained plan count is the number of distinct {@code (weight, batch shape)} keys {@code - * TornadoGgufBatchedMatrixKernel} would create for this graph. - *
- * - *

Device global-memory figures are driver-reported values, which sit below the nameplate - * capacity: the A40-4Q gate reported 3,917.7 MiB on a 4,096 MiB profile. - */ -class LargeModelEligibilityTest { - - private static final long MIB = 1024L * 1024L; - private static final long GIB = 1024L * MIB; - - /** Gemma 4 26B-A4B IT Q4_K_M: resident tensors plus routed experts. */ - private static final long GEMMA4_WEIGHT_BYTES = 1_650_027_640L + 15_130_165_248L; - - /** - * Distinct retained plans: 30 layers x (128 experts x 2 tensors + 4 resident projections) x 2 - * retained batch shapes. - */ - private static final int GEMMA4_PLAN_COUNT = 30 * (128 * 2 + 4) * 2; - - /** Device activation, scale, and output buffers summed over every retained plan. */ - private static final long GEMMA4_PLAN_SCRATCH_BYTES = 2_733_284_400L; - - /** Largest single tensor: the 262,144 x 2,816 tied embedding at Q8_0. */ - private static final long GEMMA4_LARGEST_ALLOCATION_BYTES = 262_144L * 2_816L / 32L * 34L; - - private static final long QWEN3_06B_Q4_0_BYTES = 428_970_080L; - - /** KV bytes the {@code LayeredKvCache} would hold at a runtime context length. */ - private static long gemma4KvBytes(int contextLength) { - long slidingRing = 25L * 2L * 2_048L * 4L * (1_024L + 1L); - long full = 5L * 2L * 1_024L * 4L * Math.min(contextLength, 262_144L); - return slidingRing + full; - } - - private static DeviceMemoryRequest gemma4(int contextLength, int planCount, long planScratch) { - return DeviceMemoryRequest.detailed("gemma-4-26B-A4B-it-Q4_K_M.gguf") - .weightBytes(GEMMA4_WEIGHT_BYTES) - .retainedShapes(2) - .retainedPlanCount(planCount) - .planScratchBytes(planScratch) - .largestAllocationBytes(GEMMA4_LARGEST_ALLOCATION_BYTES) - .deviceKvCacheBytes(gemma4KvBytes(contextLength)) - .build(); - } - - private static AcceleratorEligibility.DeviceCapabilities gpu(String name, long globalBytes) { - return new AcceleratorEligibility.DeviceCapabilities(name, "PTX", "GPU", globalBytes, 8L * GIB); - } - - static Stream deviceCapacityTable() { - // device, driver-reported global memory, context length, expected outcome marker - return Stream.of( - Arguments.of("A40 24 GB", 23L * GIB, 4_096, "short by"), - Arguments.of("A40 24 GB", 23L * GIB, 32_768, "short by"), - Arguments.of("A40 24 GB", 23L * GIB, 262_144, "short by"), - Arguments.of("L40S 48 GB", 44L * GIB, 4_096, "SHARED_WEIGHT_UPLOAD"), - Arguments.of("L40S 48 GB", 44L * GIB, 32_768, "SHARED_WEIGHT_UPLOAD"), - Arguments.of("L40S 48 GB", 44L * GIB, 262_144, "SHARED_WEIGHT_UPLOAD"), - Arguments.of("H100 80 GB", 79L * GIB, 4_096, "eager readiness"), - Arguments.of("H100 80 GB", 79L * GIB, 262_144, "eager readiness")); - } - - @ParameterizedTest(name = "{0} at context {2} is refused because of \"{3}\"") - @MethodSource("deviceCapacityTable") - void refusesTheShippedPlanShapeForA27bClassModelOnEveryDeviceSize( - String deviceName, long globalBytes, int contextLength, String expectedReason) { - AcceleratorEligibility.Decision decision = - AcceleratorEligibility.select( - List.of(gpu(deviceName, globalBytes)), - gemma4(contextLength, GEMMA4_PLAN_COUNT, GEMMA4_PLAN_SCRATCH_BYTES)); - - assertThat(decision.eligible()).isFalse(); - assertThat(decision.reason()).contains(expectedReason); - assertThat(decision.reason()).contains(deviceName); - } - - @Test - void itemisesEveryTermOfTheRefusedBudget() { - DeviceMemoryRequest request = gemma4(4_096, GEMMA4_PLAN_COUNT, GEMMA4_PLAN_SCRATCH_BYTES); - DeviceBudget budget = - AcceleratorEligibility.budget(request, PlanShapeStrategy.PER_SHAPE_WHOLE_MODEL); - - assertThat(budget.deviceWeightBytes()).isEqualTo(2 * GEMMA4_WEIGHT_BYTES); - assertThat(budget.kvCacheBytes()).isEqualTo(587_612_160L); - assertThat(budget.planScratchBytes()).isEqualTo(GEMMA4_PLAN_SCRATCH_BYTES); - assertThat(budget.baseOverheadBytes()).isEqualTo(256 * MIB); - assertThat(budget.totalBytes()).isEqualTo(37_149_717_792L); - assertThat(budget.hostWeightCopyBytes()).isEqualTo(2 * GEMMA4_WEIGHT_BYTES); - assertThat(budget.estimatedPlanCompileTime().toSeconds()).isEqualTo(948L); - } - - @Test - void sharedWeightUploadHalvesTheWeightTermButIsNotImplemented() { - DeviceMemoryRequest request = gemma4(4_096, GEMMA4_PLAN_COUNT, GEMMA4_PLAN_SCRATCH_BYTES); - DeviceBudget shared = - AcceleratorEligibility.budget(request, PlanShapeStrategy.SHARED_WEIGHT_UPLOAD); - - assertThat(shared.deviceWeightBytes()).isEqualTo(GEMMA4_WEIGHT_BYTES); - assertThat(shared.totalBytes()).isEqualTo(20_369_524_904L); - assertThat(PlanShapeStrategy.SHARED_WEIGHT_UPLOAD.implemented()).isFalse(); - assertThat(PlanShapeStrategy.SHARED_WEIGHT_UPLOAD.limitation()).contains("ByteArray"); - assertThat(PlanShapeStrategy.PER_SHAPE_WHOLE_MODEL.implemented()).isTrue(); - assertThat(PlanShapeStrategy.PER_SHAPE_WHOLE_MODEL.limitation()).isNull(); - assertThat(PlanShapeStrategy.BOUNDED_RESIDENT_WORKING_SET.capacityComputable()).isFalse(); - } - - @Test - void admitsAnEightyGigabyteDeviceOnceThePerExpertPlansAreGone() { - int residentOnlyPlans = 30 * 4 * 2; - AcceleratorEligibility.Decision decision = - AcceleratorEligibility.select( - List.of(gpu("H100 80 GB", 79L * GIB)), gemma4(4_096, residentOnlyPlans, 90_302_490L)); - - assertThat(decision.eligible()).isTrue(); - assertThat(decision.strategy()).isEqualTo(PlanShapeStrategy.PER_SHAPE_WHOLE_MODEL); - assertThat(decision.budget().estimatedPlanCompileTime().toSeconds()).isEqualTo(14L); - } - - @Test - void refusesAFileSizeOnlyBudgetForALargeModel() { - AcceleratorEligibility.Decision decision = - AcceleratorEligibility.select( - List.of(gpu("H100 80 GB", 79L * GIB)), - DeviceMemoryRequest.ofModelFile( - "gemma-4-26B-A4B-it-Q4_K_M.gguf", 16_796_015_136L, true)); - - assertThat(decision.eligible()).isFalse(); - assertThat(decision.reason()).contains("file-size-only device budget"); - assertThat(decision.reason()).contains("15.64 GiB of weights"); - assertThat(decision.reason()).contains("8.00 GiB"); - assertThat(decision.reason()).contains("DeviceMemoryRequest.detailed"); - } - - @Test - void refusesATensorThatCannotFitInATornadoByteArray() { - AcceleratorEligibility.Decision decision = - AcceleratorEligibility.select( - List.of(gpu("H100 80 GB", 79L * GIB)), - DeviceMemoryRequest.detailed("oversized.gguf") - .weightBytes(10L * GIB) - .retainedShapes(1) - .retainedPlanCount(8) - .planScratchBytes(MIB) - .largestAllocationBytes(3L * GIB) - .build()); - - assertThat(decision.eligible()).isFalse(); - assertThat(decision.reason()).contains("largest device allocation is 3.00 GiB"); - assertThat(decision.reason()).contains("2.00 GiB"); - assertThat(decision.reason()).contains("must be split"); - } - - @Test - void refusesADeviceWhoseSingleAllocationLimitIsTooSmall() { - AcceleratorEligibility.Decision decision = - AcceleratorEligibility.select( - List.of( - new AcceleratorEligibility.DeviceCapabilities( - "A40-4Q", "PTX", "GPU", 79L * GIB, 512L * MIB)), - gemma4(4_096, 240, 90_302_490L)); - - assertThat(decision.eligible()).isFalse(); - assertThat(decision.reason()).contains("single allocation of at most 0.50 GiB"); - assertThat(decision.reason()).contains("needs one of 0.73 GiB"); - } - - @Test - void namesEveryDiscoveredDeviceWhenNoneQualifies() { - AcceleratorEligibility.Decision decision = - AcceleratorEligibility.select( - List.of( - new AcceleratorEligibility.DeviceCapabilities( - "AMD Radeon", "OPENCL", "GPU", 24L * GIB, 4L * GIB), - new AcceleratorEligibility.DeviceCapabilities( - "host CPU", "JAVA", "CPU", 64L * GIB, 16L * GIB)), - gemma4(4_096, 240, 90_302_490L)); - - assertThat(decision.eligible()).isFalse(); - assertThat(decision.reason()).contains("PTX or CUDA backend on a GPU device"); - assertThat(decision.reason()).contains("AMD Radeon [OPENCL/GPU]"); - assertThat(decision.reason()).contains("host CPU [JAVA/CPU]"); - } - - @Test - void reportsNoDevicesWhenTheRuntimeFoundNone() { - AcceleratorEligibility.Decision decision = - AcceleratorEligibility.select(List.of(), gemma4(4_096, 240, 90_302_490L)); - - assertThat(decision.eligible()).isFalse(); - assertThat(decision.reason()).contains("no devices"); - } - - @Test - void keepsThePublishedCoarseBudgetForTheQualifiedSmallModel() { - DeviceMemoryRequest request = - DeviceMemoryRequest.ofModelFile("Qwen3-0.6B-Q4_0.gguf", QWEN3_06B_Q4_0_BYTES, true); - DeviceBudget budget = - AcceleratorEligibility.budget(request, PlanShapeStrategy.PER_SHAPE_WHOLE_MODEL); - - assertThat(request.detailed()).isFalse(); - assertThat(budget.totalBytes()).isEqualTo(1_126_375_616L); - assertThat(budget.kvCacheBytes()).isZero(); - assertThat(budget.estimatedPlanCompileTime()).isZero(); - } - - @Test - void carriesADeviceResidentKvCacheOntoAnExistingRequest() { - DeviceMemoryRequest base = - DeviceMemoryRequest.ofModelFile("Qwen3-0.6B-Q4_0.gguf", QWEN3_06B_Q4_0_BYTES, false); - - DeviceMemoryRequest withKv = base.withDeviceKvCacheBytes(64 * MIB); - - assertThat(base.deviceKvCacheBytes()).isZero(); - assertThat(withKv.deviceKvCacheBytes()).isEqualTo(64 * MIB); - assertThat(withKv.retainedShapes()).isEqualTo(1); - assertThat( - AcceleratorEligibility.budget(withKv, PlanShapeStrategy.PER_SHAPE_WHOLE_MODEL) - .totalBytes()) - .isEqualTo(QWEN3_06B_Q4_0_BYTES + 64 * MIB + 256 * MIB); - } - - @Test - void rejectsMalformedRequests() { - assertThat( - org.assertj.core.api.Assertions.catchThrowable( - () -> DeviceMemoryRequest.ofModelFile(" ", 1, true))) - .isInstanceOf(IllegalArgumentException.class); - assertThat( - org.assertj.core.api.Assertions.catchThrowable( - () -> DeviceMemoryRequest.detailed("m").weightBytes(0).build())) - .isInstanceOf(IllegalArgumentException.class); - assertThat( - org.assertj.core.api.Assertions.catchThrowable( - () -> DeviceMemoryRequest.detailed("m").weightBytes(1).retainedShapes(0).build())) - .isInstanceOf(IllegalArgumentException.class); - assertThat( - org.assertj.core.api.Assertions.catchThrowable( - () -> - DeviceMemoryRequest.detailed("m") - .weightBytes(1) - .deviceKvCacheBytes(-1) - .build())) - .isInstanceOf(IllegalArgumentException.class); - assertThat( - org.assertj.core.api.Assertions.catchThrowable( - () -> - PlanShapeStrategy.BOUNDED_RESIDENT_WORKING_SET.deviceWeightBytes( - DeviceMemoryRequest.ofModelFile("m", 1, false)))) - .isInstanceOf(IllegalStateException.class); - } -} diff --git a/backend-tornado/src/test/java/com/integrallis/models/backend/tornado/Q4ProjectionKernelTest.java b/backend-tornado/src/test/java/com/integrallis/models/backend/tornado/Q4ProjectionKernelTest.java deleted file mode 100644 index dff73fc1..00000000 --- a/backend-tornado/src/test/java/com/integrallis/models/backend/tornado/Q4ProjectionKernelTest.java +++ /dev/null @@ -1,200 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.backend.tornado; - -import static org.assertj.core.api.Assertions.assertThat; -import static org.assertj.core.api.Assertions.assertThatThrownBy; -import static org.assertj.core.api.Assertions.within; - -import com.integrallis.vectors.core.VectorUtil; -import java.lang.foreign.MemorySegment; -import java.util.Random; -import org.junit.jupiter.api.Test; -import uk.ac.manchester.tornado.api.types.arrays.ByteArray; -import uk.ac.manchester.tornado.api.types.arrays.FloatArray; - -class Q4ProjectionKernelTest { - - @Test - void matchesProductionQ4ByQ8ProjectionAcrossBatches() { - int batchSize = 3; - int rows = 17; - int cols = 96; - byte[] weights = randomQ4Matrix(rows, cols, 31L); - float[] input = randomFloats(batchSize * cols, 37L); - byte[] activations = new byte[batchSize * cols]; - float[] activationScales = new float[batchSize * cols / 32]; - Q4ProjectionKernel.quantize(input, activations, activationScales, batchSize, cols); - float[] expected = vectorApiProjection(weights, input, batchSize, rows, cols); - ByteArray deviceWeights = ByteArray.fromArray(weights); - ByteArray deviceActivations = ByteArray.fromArray(activations); - FloatArray deviceScales = FloatArray.fromArray(activationScales); - FloatArray deviceOutput = new FloatArray(batchSize * rows); - - Q4ProjectionKernel.multiply( - deviceWeights, deviceActivations, deviceScales, deviceOutput, batchSize, rows, cols); - - float[] actual = deviceOutput.toHeapArray(); - assertThat(actual).hasSameSizeAs(expected); - for (int index = 0; index < expected.length; index++) { - assertThat(actual[index]).isCloseTo(expected[index], within(2.0e-5f)); - } - } - - @Test - void matchesTwoProductionProjectionsWithOnePreparedActivation() { - int batchSize = 3; - int firstRows = 11; - int secondRows = 7; - int cols = 96; - byte[] firstWeights = randomQ4Matrix(firstRows, cols, 41L); - byte[] secondWeights = randomQ4Matrix(secondRows, cols, 43L); - float[] input = randomFloats(batchSize * cols, 47L); - byte[] activations = new byte[batchSize * cols]; - float[] scales = new float[batchSize * cols / 32]; - Q4ProjectionKernel.quantize(input, activations, scales, batchSize, cols); - float[] expectedFirst = vectorApiProjection(firstWeights, input, batchSize, firstRows, cols); - float[] expectedSecond = vectorApiProjection(secondWeights, input, batchSize, secondRows, cols); - FloatArray actualFirst = new FloatArray(batchSize * firstRows); - FloatArray actualSecond = new FloatArray(batchSize * secondRows); - - Q4ProjectionKernel.multiplyDual( - ByteArray.fromArray(firstWeights), - firstRows, - ByteArray.fromArray(secondWeights), - secondRows, - ByteArray.fromArray(activations), - FloatArray.fromArray(scales), - actualFirst, - actualSecond, - batchSize, - cols); - - assertClose(actualFirst.toHeapArray(), expectedFirst); - assertClose(actualSecond.toHeapArray(), expectedSecond); - } - - @Test - void matchesThreeProductionProjectionsWithOnePreparedActivation() { - int batchSize = 2; - int firstRows = 9; - int secondRows = 7; - int thirdRows = 5; - int cols = 64; - byte[] firstWeights = randomQ4Matrix(firstRows, cols, 53L); - byte[] secondWeights = randomQ4Matrix(secondRows, cols, 59L); - byte[] thirdWeights = randomQ4Matrix(thirdRows, cols, 61L); - float[] input = randomFloats(batchSize * cols, 67L); - byte[] activations = new byte[batchSize * cols]; - float[] scales = new float[batchSize * cols / 32]; - Q4ProjectionKernel.quantize(input, activations, scales, batchSize, cols); - float[] expectedFirst = vectorApiProjection(firstWeights, input, batchSize, firstRows, cols); - float[] expectedSecond = vectorApiProjection(secondWeights, input, batchSize, secondRows, cols); - float[] expectedThird = vectorApiProjection(thirdWeights, input, batchSize, thirdRows, cols); - FloatArray actualFirst = new FloatArray(batchSize * firstRows); - FloatArray actualSecond = new FloatArray(batchSize * secondRows); - FloatArray actualThird = new FloatArray(batchSize * thirdRows); - - Q4ProjectionKernel.multiplyTriple( - ByteArray.fromArray(firstWeights), - firstRows, - ByteArray.fromArray(secondWeights), - secondRows, - ByteArray.fromArray(thirdWeights), - thirdRows, - ByteArray.fromArray(activations), - FloatArray.fromArray(scales), - actualFirst, - actualSecond, - actualThird, - batchSize, - cols); - - assertClose(actualFirst.toHeapArray(), expectedFirst); - assertClose(actualSecond.toHeapArray(), expectedSecond); - assertClose(actualThird.toHeapArray(), expectedThird); - } - - @Test - void rejectsMismatchedActivationScaleStorage() { - assertThatThrownBy( - () -> - Q4ProjectionKernel.validate( - new ByteArray(18), - new ByteArray(32), - new FloatArray(0), - new FloatArray(1), - 1, - 1, - 32)) - .isInstanceOf(IllegalArgumentException.class) - .hasMessageContaining("scales"); - } - - private static float[] vectorApiProjection( - byte[] weights, float[] input, int batchSize, int rows, int cols) { - float[] output = new float[batchSize * rows]; - MemorySegment weightSegment = MemorySegment.ofArray(weights); - for (int batch = 0; batch < batchSize; batch++) { - float[] query = new float[cols]; - float[] projected = new float[rows]; - System.arraycopy(input, batch * cols, query, 0, cols); - VectorUtil.ggufQ4_0Q8_0BatchDotProduct( - query, - weightSegment, - rows, - cols, - projected, - new byte[cols], - new float[cols / 32], - new int[(cols + 3) / 4]); - System.arraycopy(projected, 0, output, batch * rows, rows); - } - return output; - } - - private static void assertClose(float[] actual, float[] expected) { - assertThat(actual).hasSameSizeAs(expected); - for (int index = 0; index < expected.length; index++) { - assertThat(actual[index]).isCloseTo(expected[index], within(2.0e-5f)); - } - } - - private static byte[] randomQ4Matrix(int rows, int cols, long seed) { - Random random = new Random(seed); - int blocks = rows * cols / 32; - byte[] weights = new byte[blocks * 18]; - for (int block = 0; block < blocks; block++) { - short scale = Float.floatToFloat16(0.001f + random.nextFloat() * 0.05f); - int offset = block * 18; - weights[offset] = (byte) scale; - weights[offset + 1] = (byte) (scale >>> 8); - for (int quant = 0; quant < 16; quant++) { - weights[offset + 2 + quant] = (byte) random.nextInt(256); - } - } - return weights; - } - - private static float[] randomFloats(int length, long seed) { - Random random = new Random(seed); - float[] values = new float[length]; - for (int index = 0; index < length; index++) { - values[index] = random.nextFloat(-2.0f, 2.0f); - } - return values; - } -} diff --git a/backend-tornado/src/test/java/com/integrallis/models/backend/tornado/TornadoBackendIntegrationTest.java b/backend-tornado/src/test/java/com/integrallis/models/backend/tornado/TornadoBackendIntegrationTest.java deleted file mode 100644 index e8884ecb..00000000 --- a/backend-tornado/src/test/java/com/integrallis/models/backend/tornado/TornadoBackendIntegrationTest.java +++ /dev/null @@ -1,59 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.backend.tornado; - -import static org.assertj.core.api.Assertions.assertThat; -import static org.junit.jupiter.api.Assumptions.assumeTrue; - -import java.nio.file.Files; -import java.nio.file.Path; -import org.junit.jupiter.api.Tag; -import org.junit.jupiter.api.Test; - -@Tag("integration") -class TornadoBackendIntegrationTest { - - @Test - void loadsAndRunsThePinnedQwenModelThroughAutomaticSelectionWhenProvided() { - String configured = System.getProperty("models.fixtures.qwen306BQ40", ""); - assumeTrue(!configured.isBlank(), "set -Dmodels.fixtures.qwen306BQ40="); - Path model = Path.of(configured).toAbsolutePath().normalize(); - assumeTrue(Files.isRegularFile(model), "Qwen fixture is not installed"); - boolean required = Boolean.getBoolean("models.accelerator.required"); - TornadoBackendOptions options = new TornadoBackendOptions(true, true, required, 32); - - try (TornadoBackendRuntime runtime = TornadoBackend.open(model, options)) { - int token = runtime.backend().tokenizer().bosToken(); - if (token < 0 || token >= runtime.backend().metadata().vocabSize()) { - token = 0; - } - float[] logits = runtime.backend().prefill(new int[] {token}, 0); - - assertThat(logits).hasSize(runtime.backend().metadata().vocabSize()); - for (float logit : logits) { - assertThat(Float.isFinite(logit)).isTrue(); - } - if (required) { - assertThat(runtime.status().accelerated()).isTrue(); - assertThat(runtime.status().readinessTime()).isPositive(); - } - String expected = System.getProperty("models.accelerator.expected", ""); - if (!expected.isBlank()) { - assertThat(runtime.status().accelerated()).isEqualTo(Boolean.parseBoolean(expected)); - } - } - } -} diff --git a/backend-tornado/src/test/java/com/integrallis/models/backend/tornado/TornadoBackendOptionsTest.java b/backend-tornado/src/test/java/com/integrallis/models/backend/tornado/TornadoBackendOptionsTest.java deleted file mode 100644 index 8146700c..00000000 --- a/backend-tornado/src/test/java/com/integrallis/models/backend/tornado/TornadoBackendOptionsTest.java +++ /dev/null @@ -1,40 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.backend.tornado; - -import static org.assertj.core.api.Assertions.assertThat; - -import org.junit.jupiter.api.Test; - -class TornadoBackendOptionsTest { - - @Test - void defaultsPrepareBothPrefillAndDecodeAndAllowCpuFallback() { - TornadoBackendOptions options = TornadoBackendOptions.defaults(); - - assertThat(options.accelerateDecode()).isTrue(); - assertThat(options.eagerReadiness()).isTrue(); - assertThat(options.requireAccelerator()).isFalse(); - assertThat(options.executionBatchSize()).isEqualTo(32); - } - - @Test - void validatesTheFixedDeviceBatchShape() { - org.assertj.core.api.Assertions.assertThatIllegalArgumentException() - .isThrownBy(() -> new TornadoBackendOptions(true, true, false, 3)) - .withMessageContaining("executionBatchSize"); - } -} diff --git a/backend-tornado/src/test/java/com/integrallis/models/backend/tornado/TornadoBackendStatusTest.java b/backend-tornado/src/test/java/com/integrallis/models/backend/tornado/TornadoBackendStatusTest.java deleted file mode 100644 index 6072cb0b..00000000 --- a/backend-tornado/src/test/java/com/integrallis/models/backend/tornado/TornadoBackendStatusTest.java +++ /dev/null @@ -1,55 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.backend.tornado; - -import static org.assertj.core.api.Assertions.assertThat; -import static org.assertj.core.api.Assertions.assertThatThrownBy; - -import java.time.Duration; -import org.junit.jupiter.api.Test; - -class TornadoBackendStatusTest { - - @Test - void normalizesAValidStatus() { - TornadoBackendStatus status = - new TornadoBackendStatus( - true, " NVIDIA A40 ", " eligible ", 1024, Duration.ofSeconds(2)); - - assertThat(status.device()).isEqualTo("NVIDIA A40"); - assertThat(status.reason()).isEqualTo("eligible"); - assertThat(status.requiredDeviceBytes()).isEqualTo(1024); - assertThat(status.readinessTime()).isEqualTo(Duration.ofSeconds(2)); - } - - @Test - void rejectsInvalidStatusValues() { - assertThatThrownBy(() -> new TornadoBackendStatus(false, " ", "unavailable", 0, Duration.ZERO)) - .isInstanceOf(IllegalArgumentException.class) - .hasMessageContaining("device"); - assertThatThrownBy(() -> new TornadoBackendStatus(false, "CPU", " ", 0, Duration.ZERO)) - .isInstanceOf(IllegalArgumentException.class) - .hasMessageContaining("reason"); - assertThatThrownBy( - () -> new TornadoBackendStatus(false, "CPU", "unavailable", -1, Duration.ZERO)) - .isInstanceOf(IllegalArgumentException.class) - .hasMessageContaining("requiredDeviceBytes"); - assertThatThrownBy( - () -> new TornadoBackendStatus(false, "CPU", "unavailable", 0, Duration.ofNanos(-1))) - .isInstanceOf(IllegalArgumentException.class) - .hasMessageContaining("readinessTime"); - } -} diff --git a/backend-tornado/src/test/java/com/integrallis/models/backend/tornado/TornadoBackendTest.java b/backend-tornado/src/test/java/com/integrallis/models/backend/tornado/TornadoBackendTest.java deleted file mode 100644 index f7b185a1..00000000 --- a/backend-tornado/src/test/java/com/integrallis/models/backend/tornado/TornadoBackendTest.java +++ /dev/null @@ -1,50 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.backend.tornado; - -import static org.assertj.core.api.Assertions.assertThatIllegalArgumentException; -import static org.assertj.core.api.Assertions.assertThatNullPointerException; - -import com.integrallis.models.api.BackendConfiguration; -import java.nio.file.Files; -import java.nio.file.Path; -import org.junit.jupiter.api.Test; -import org.junit.jupiter.api.io.TempDir; - -class TornadoBackendTest { - - @Test - void rejectsAMissingModelBeforeInspectingAcceleratorDrivers(@TempDir Path directory) { - Path missing = directory.resolve("missing.gguf"); - - assertThatIllegalArgumentException() - .isThrownBy(() -> TornadoBackend.open(missing)) - .withMessageContaining("regular model file"); - } - - @Test - void validatesConfigurationBeforeInspectingAcceleratorDrivers(@TempDir Path directory) - throws Exception { - Path placeholder = Files.createFile(directory.resolve("model.gguf")); - - assertThatNullPointerException() - .isThrownBy(() -> TornadoBackend.open(placeholder, null, TornadoBackendOptions.defaults())) - .withMessage("backendConfiguration"); - assertThatNullPointerException() - .isThrownBy(() -> TornadoBackend.open(placeholder, BackendConfiguration.empty(), null)) - .withMessage("options"); - } -} diff --git a/backend-tornado/src/test/java/com/integrallis/models/backend/tornado/TornadoGgufBatchedMatrixKernelTest.java b/backend-tornado/src/test/java/com/integrallis/models/backend/tornado/TornadoGgufBatchedMatrixKernelTest.java deleted file mode 100644 index 2c4d3ecf..00000000 --- a/backend-tornado/src/test/java/com/integrallis/models/backend/tornado/TornadoGgufBatchedMatrixKernelTest.java +++ /dev/null @@ -1,210 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.backend.tornado; - -import static org.assertj.core.api.Assertions.assertThat; -import static org.assertj.core.api.Assertions.assertThatThrownBy; - -import com.integrallis.models.backend.purejava.gguf.GgufTensorType; -import com.integrallis.models.backend.purejava.plan.PureJavaPlanConfiguration; -import java.lang.foreign.MemorySegment; -import org.junit.jupiter.api.Test; - -class TornadoGgufBatchedMatrixKernelTest { - - @Test - void acceleratesQ4PrefillButLeavesDecodeAndUnsupportedFormatsOnTheJavaPath() { - try (TornadoGgufBatchedMatrixKernel kernel = new TornadoGgufBatchedMatrixKernel()) { - assertThat(kernel.executionBatchSize()).isEqualTo(32); - assertThat(kernel.supports(GgufTensorType.Q4_0)).isTrue(); - assertThat(kernel.supports(GgufTensorType.Q4_K)).isTrue(); - assertThat(kernel.supports(GgufTensorType.Q6_K)).isTrue(); - assertThat(kernel.supports(GgufTensorType.Q5_K)).isFalse(); - assertThat(kernel.supports(GgufTensorType.Q8_0)).isFalse(); - assertThat(kernel.supports(GgufTensorType.F32)).isFalse(); - assertThat(kernel.isEligible(GgufTensorType.Q4_0, 1, 3072, 1024)).isFalse(); - assertThat(kernel.isEligible(GgufTensorType.Q4_0, 4, 3072, 1024)).isTrue(); - assertThat(kernel.isEligible(GgufTensorType.Q4_0, 30, 3072, 1024)).isTrue(); - assertThat(kernel.isEligible(GgufTensorType.Q4_0, 32, 3072, 1024)).isTrue(); - assertThat(kernel.isEligible(GgufTensorType.Q4_0, 33, 3072, 1024)).isFalse(); - assertThat(kernel.isEligible(GgufTensorType.Q4_0, 4, 32, 32)).isFalse(); - assertThat(kernel.supportsDual(GgufTensorType.Q4_0, GgufTensorType.Q4_0)).isTrue(); - assertThat( - kernel.supportsTriple(GgufTensorType.Q4_0, GgufTensorType.Q4_0, GgufTensorType.Q4_0)) - .isTrue(); - } - } - - @Test - void acceptsAnExplicitFixedExecutionBatch() { - try (TornadoGgufBatchedMatrixKernel kernel = new TornadoGgufBatchedMatrixKernel(64)) { - assertThat(kernel.executionBatchSize()).isEqualTo(64); - assertThat(kernel.isEligible(GgufTensorType.Q4_0, 64, 3072, 1024)).isTrue(); - assertThat(kernel.isEligible(GgufTensorType.Q4_0, 65, 3072, 1024)).isFalse(); - } - } - - @Test - void decodeAccelerationIsExplicitAndUsesItsOwnSingleTokenShape() { - try (TornadoGgufBatchedMatrixKernel kernel = new TornadoGgufBatchedMatrixKernel(32, true)) { - assertThat(kernel.acceleratesDecode()).isTrue(); - assertThat(kernel.executionBatchSizeFor(1)).isEqualTo(1); - assertThat(kernel.executionBatchSizeFor(4)).isEqualTo(32); - assertThat(kernel.isEligible(GgufTensorType.Q4_0, 1, 3072, 1024)).isTrue(); - assertThat( - kernel.isDualEligible(GgufTensorType.Q4_0, 1024, GgufTensorType.Q4_0, 1024, 1, 1024)) - .isTrue(); - } - } - - @Test - void rejectsAnExecutionBatchBelowTheGpuThreshold() { - assertThatThrownBy(() -> new TornadoGgufBatchedMatrixKernel(3)) - .isInstanceOf(IllegalArgumentException.class) - .hasMessageContaining("executionBatchSize"); - } - - @Test - void padsTheLastPromptChunkWithoutReusingStaleActivations() { - float[] padded = {9.0f, 9.0f, 9.0f, 9.0f, 9.0f, 9.0f, 9.0f, 9.0f}; - - float[] executionInput = - TornadoGgufBatchedMatrixKernel.prepareExecutionInput( - new float[] {1.0f, 2.0f, 3.0f, 4.0f}, padded, 2, 4, 2); - - assertThat(executionInput).isSameAs(padded); - assertThat(executionInput).containsExactly(1.0f, 2.0f, 3.0f, 4.0f, 0.0f, 0.0f, 0.0f, 0.0f); - } - - @Test - void recommendsTheGroupedProjectionPlanThatTheExperimentImplements() { - try (TornadoGgufBatchedMatrixKernel kernel = new TornadoGgufBatchedMatrixKernel()) { - assertThat(kernel.planRecommendations()) - .containsEntry(PureJavaPlanConfiguration.GROUPED_PROJECTIONS_PROPERTY, "true") - .containsEntry(PureJavaPlanConfiguration.STAGED_QUANTIZED_FFN_PROPERTY, "false") - .containsEntry(PureJavaPlanConfiguration.STAGED_QUANTIZED_LAYER_PROPERTY, "false"); - } - } - - @Test - void admitsTheMixedKProjectionGroupThatQ4KMModelsPresent() { - try (TornadoGgufBatchedMatrixKernel kernel = new TornadoGgufBatchedMatrixKernel()) { - assertThat(kernel.planRecommendations()) - .containsEntry(PureJavaPlanConfiguration.MIXED_K_PROJECTIONS_PROPERTY, "true"); - assertThat( - kernel.supportsTriple(GgufTensorType.Q4_K, GgufTensorType.Q4_K, GgufTensorType.Q6_K)) - .isTrue(); - assertThat( - kernel.isTripleEligible( - GgufTensorType.Q4_K, - 2048, - GgufTensorType.Q4_K, - 512, - GgufTensorType.Q6_K, - 512, - 8, - 2048)) - .isTrue(); - } - } - - @Test - void keepsTheTwoActivationFamiliesOutOfOneGroupedDispatch() { - try (TornadoGgufBatchedMatrixKernel kernel = new TornadoGgufBatchedMatrixKernel()) { - // Q4_0 needs Q8_0 activations and the K-quants need Q8_K, so one prepared activation can - // never serve both. Mixed groups must fall back rather than silently use the wrong scales. - assertThat(kernel.supportsDual(GgufTensorType.Q4_0, GgufTensorType.Q4_K)).isFalse(); - assertThat(kernel.supportsDual(GgufTensorType.Q4_K, GgufTensorType.Q4_0)).isFalse(); - assertThat( - kernel.supportsTriple(GgufTensorType.Q4_K, GgufTensorType.Q4_0, GgufTensorType.Q4_K)) - .isFalse(); - // Q6_K has no dual kernel and is only ever the third matrix of a grouped dispatch. - assertThat(kernel.supportsDual(GgufTensorType.Q6_K, GgufTensorType.Q6_K)).isFalse(); - assertThat( - kernel.supportsTriple(GgufTensorType.Q6_K, GgufTensorType.Q4_K, GgufTensorType.Q4_K)) - .isFalse(); - assertThat(kernel.supportsDual(GgufTensorType.Q4_K, GgufTensorType.Q4_K)).isTrue(); - assertThat(kernel.supportsDual(GgufTensorType.Q5_K, GgufTensorType.Q5_K)).isFalse(); - } - } - - @Test - void requiresWholeSuperBlocksForKQuantProjections() { - try (TornadoGgufBatchedMatrixKernel kernel = new TornadoGgufBatchedMatrixKernel()) { - assertThat(kernel.isEligible(GgufTensorType.Q4_K, 8, 4096, 1024)).isTrue(); - assertThat(kernel.isEligible(GgufTensorType.Q6_K, 8, 4096, 1024)).isTrue(); - // 1120 is a multiple of 32 but not of 256: legal for Q4_0, never for a K-quant. - assertThat(kernel.isEligible(GgufTensorType.Q4_0, 8, 4096, 1120)).isTrue(); - assertThat(kernel.isEligible(GgufTensorType.Q4_K, 8, 4096, 1120)).isFalse(); - assertThat( - kernel.isDualEligible(GgufTensorType.Q4_K, 4096, GgufTensorType.Q4_K, 4096, 8, 1120)) - .isFalse(); - } - } - - @Test - void refusesTensorsTooLargeForAnIntIndexedDeviceBuffer() { - try (TornadoGgufBatchedMatrixKernel kernel = new TornadoGgufBatchedMatrixKernel()) { - // Q4_K stores 144 bytes per 256 values, so a tensor needs about 3.8e9 values before it - // stops fitting an int-indexed device buffer. A 27B-class vocabulary projection - // (262144 x 5120 = 755 MiB of Q4_K) is comfortably inside that; a tensor four times - // taller is not, and must stay on the Vector API rather than fail the load. - assertThat(kernel.isEligible(GgufTensorType.Q4_K, 8, 262_144, 5120)).isTrue(); - assertThat(kernel.isEligible(GgufTensorType.Q4_K, 8, 1_000_000, 5120)).isFalse(); - assertThat(kernel.isEligible(GgufTensorType.Q6_K, 8, 1_000_000, 5120)).isFalse(); - assertThat(kernel.isEligible(GgufTensorType.Q4_0, 8, 4_000_000, 5120)).isFalse(); - assertThat( - kernel.isTripleEligible( - GgufTensorType.Q4_K, - 1_000_000, - GgufTensorType.Q4_K, - 512, - GgufTensorType.Q6_K, - 512, - 8, - 5120)) - .isFalse(); - } - } - - @Test - void reportsNoRoutedProjectionsBeforeAnyDispatch() { - try (TornadoGgufBatchedMatrixKernel kernel = new TornadoGgufBatchedMatrixKernel()) { - assertThat(kernel.routedProjectionsByFormat()).isEmpty(); - assertThat(kernel.projectionPlanCount()).isZero(); - assertThat(kernel.calls()).isZero(); - assertThat(kernel.totalMillis()).isZero(); - } - } - - @Test - void refusesToDispatchProjectionsItDeclaredIneligible() { - try (TornadoGgufBatchedMatrixKernel kernel = new TornadoGgufBatchedMatrixKernel()) { - assertThatThrownBy( - () -> - kernel.multiply( - new float[8], - new float[8], - MemorySegment.ofArray(new byte[8]), - GgufTensorType.Q5_K, - 8, - 1, - 1)) - .isInstanceOf(UnsupportedOperationException.class) - .hasMessageContaining("not eligible"); - } - } -} diff --git a/models-accelerator-bench/results/rtx4090-ptx-2026-10-07.md b/benchmark-results/2026-10-07-tornado-removal/README.md similarity index 77% rename from models-accelerator-bench/results/rtx4090-ptx-2026-10-07.md rename to benchmark-results/2026-10-07-tornado-removal/README.md index 38812845..d7c2a918 100644 --- a/models-accelerator-bench/results/rtx4090-ptx-2026-10-07.md +++ b/benchmark-results/2026-10-07-tornado-removal/README.md @@ -1,4 +1,27 @@ -# backend-tornado on RTX 4090 / PTX: the gate returns 0 and the device errors 36 times +# Why backend-tornado was removed + +**Decision 2026-10-07: Models ships one GPU implementation.** The measurement below is why the +TornadoVM arm is not it. Same RTX 4090, same model, same day as `backend-cuda`'s gates: + +| | `backend-cuda` | `backend-tornado` | +| --- | --- | --- | +| G1 exact token parity | **1280 / 1280 identical** | not reached | +| decode | **31.26 tok/s** | 6.49 tok/s | +| CPU control, same host | 5.06 tok/s | 5.06 tok/s | +| readiness | 344 ms | 32,171 ms | +| device launch errors | 0 | **36 x CUDA 701** | +| self-contained artifact | yes | no -- needs a separately installed runtime | + +A second implementation of the same capability, five times slower than the first and barely ahead of +the CPU path it is meant to accelerate, is not worth the surface it costs. It is removed rather than +carried: `backend-tornado`, its publication-allowlist entry, the `accelerator-profile` command it +drove, its TornadoVM benchmark arm, and the release precondition that pointed at its results +directory. + +The original run notes follow, kept because the numbers above come from them and because the +`cuLaunchKernel` failures are the substantive finding. + +## The run as it was recorded **Measured 2026-10-07.** RTX 4090, compute capability 8.9. TornadoVM **v5.2.0-jdk25** built from source with the PTX backend, JDK 25. Models revision `10af16c33be7e023c114812e6137b47fab7e0c90`. diff --git a/models-accelerator-bench/results/rtx4090-ptx-2026-10-07-accelerator-profile.json b/benchmark-results/2026-10-07-tornado-removal/accelerator-profile.json similarity index 100% rename from models-accelerator-bench/results/rtx4090-ptx-2026-10-07-accelerator-profile.json rename to benchmark-results/2026-10-07-tornado-removal/accelerator-profile.json diff --git a/models-accelerator-bench/results/rtx4090-ptx-2026-10-07-worker.log b/benchmark-results/2026-10-07-tornado-removal/worker.log similarity index 100% rename from models-accelerator-bench/results/rtx4090-ptx-2026-10-07-worker.log rename to benchmark-results/2026-10-07-tornado-removal/worker.log diff --git a/build.gradle.kts b/build.gradle.kts index d3ddbcb2..33378191 100644 --- a/build.gradle.kts +++ b/build.gradle.kts @@ -51,7 +51,6 @@ val publishedModuleNames = "models-rag", "models-semantic-order", "backend-java", - "backend-tornado", "backend-native", "backend-cuda", "backend-apple", diff --git a/docs/content/modules/ROOT/pages/architecture.adoc b/docs/content/modules/ROOT/pages/architecture.adoc index dbd5e113..a05069c4 100644 --- a/docs/content/modules/ROOT/pages/architecture.adoc +++ b/docs/content/modules/ROOT/pages/architecture.adoc @@ -30,10 +30,12 @@ loaded through Java's final FFM API. It does not embed or call llama.cpp or Ollama. The Java graph remains authoritative, so the backend boundary is narrow and testable. -`backend-tornado` is a Java-authored device path for that same graph. TornadoVM -compiles eligible Q4_0 projections from Java bytecode for a qualified NVIDIA -GPU; attention and unsupported projections remain on the Vector API. No model -server or second inference engine is involved. +`backend-cuda` is the device path for that same graph. Models-owned Rust kernels +compiled to PTX run the eligible K-quant projections and grouped-query decode +attention on a qualified NVIDIA GPU; unsupported formats and shapes remain on +the Vector API, and a declined projection is counted rather than silently +dropped. No model server or second inference engine is involved, and Models +ships only this one GPU implementation. `backend-apple` is also in-process: Java reaches Apple's on-device Foundation Models framework through a bundled, integrity-checked FFM diff --git a/docs/content/modules/ROOT/pages/gpu-acceleration.adoc b/docs/content/modules/ROOT/pages/gpu-acceleration.adoc index 53f0aea4..b570f436 100644 --- a/docs/content/modules/ROOT/pages/gpu-acceleration.adoc +++ b/docs/content/modules/ROOT/pages/gpu-acceleration.adoc @@ -1,188 +1,107 @@ -= Java GPU Acceleration - -`backend-tornado` optionally executes Models-owned Java projection kernels -on a qualified NVIDIA GPU. The model parser, tokenizer, transformer graph, -attention, KV cache, sampling, and generation loop remain inside Models. There -is no external inference server and no Rust or handwritten CUDA kernel in this -path. - -== Qualified Scope - -The production selector currently admits: - -* GGUF Q4_0, Q4_K, and Q6_K projection work; -* NVIDIA devices exposed by TornadoVM's PTX backend; -* a device with enough memory for retained prefill and decode plans plus the - required safety margin; -* a single weight tensor small enough to address with a 32-bit index - (under 2 GiB on the device); and -* fixed 32-token prefill plans and separate single-token decode plans. - -Attention and unsupported tensors continue on the Vector API. AMD, Intel, and -Metal devices remain CPU fallback paths until each vendor path passes a real -hardware parity and performance gate. - -=== Weight Formats And Activation Families - -Q4_0 weights are multiplied by Q8_0 activations; Q4_K and Q6_K weights are -multiplied by Q8_K activations. A grouped dispatch shares one prepared -activation across two or three weight tensors, so the two families are never -combined inside one group even though a `Q4_K_M` model contains both. The -admitted grouped shapes are `Q4_0/Q4_0`, `Q4_0/Q4_0/Q4_0`, `Q4_K/Q4_K`, -`Q4_K/Q4_K/Q4_K`, and `Q4_K/Q4_K/Q6_K` — the last being the query, key, and -value group a `Q4_K_M` model presents, since llama.cpp promotes the value -projection to Q6_K. Any other combination falls back to the Vector API -per projection, and Q5_K, Q2_K, Q3_K, Q8_0, F32, and BF16 tensors are never -routed to the device. - -K-quant tensors additionally require a column count that is a whole multiple of -256, the K-quant super-block size. Q4_0's 32-value blocks are not sufficient. - -=== Numeric Parity With The CPU Kernels - -The K-quant kernels accumulate every per-super-block quantized dot product and -minimum correction in `int`, so those reductions are exact and independent of -how work is scheduled. Only the per-super-block scale application is floating -point, and it runs in ascending super-block order in the same two-step form the -vectors-core CPU kernels use. The CPU kernels fuse those steps with -`Math.fma` and the device kernels use a plain multiply and add, so the -guaranteed contract is agreement to within two float roundings per super-block. -Off-device parity tests assert both halves of this: bit-for-bit equality on -super-blocks whose scales are powers of two, and the rounding budget on -pseudo-random super-blocks. - -Device-side parity — that TornadoVM's PTX backend lowers these kernels to the -same arithmetic — is a hardware gate, not a unit test. - -== Large Models Are Refused, Not Guessed At - -The capacity gate adds up what a model would place on the device: weights under -the plan shape the kernel builds, per-plan scratch, and any device-resident KV -cache. It refuses a model when a term it needs is unknown rather than assuming -it is zero, so a model above 8 GiB of weights is not admitted on a file-size-only -budget. It also refuses a tensor at or above the 2 GiB TornadoVM `ByteArray` -limit, and a plan set whose eager readiness would exceed 120 seconds. - -A 27B-class Q4_K_M model does not run on this path today. It holds no Q4_0 -tensors, and the retained plan cache keeps one host and one device copy of every -weight per batch shape, so the shipped plan shape asks for roughly twice the -model on both sides. Ineligibility messages name what was needed, what was -available, and which plan shape would have fit, so a refusal is actionable -rather than a bare fallback. - -== Add The Optional Backend - -[source,kotlin,subs="attributes+"] ----- -dependencies { - implementation("com.integrallis:backend-tornado:{models-version}") -} ----- - -Install a TornadoVM distribution compatible with the application's JDK and -with its PTX backend enabled. Follow the -https://tornadovm.readthedocs.io/en/latest/installation.html[official TornadoVM installation guide^]. -The `backend-tornado` dependency provides the Models integration but does not -bundle or transitively install TornadoVM's device runtime. Device discovery -requires the TornadoVM distribution and launch configuration. Start the -application with `tornado`, or with ordinary `java` and TornadoVM's generated -argument file. += NVIDIA GPU Acceleration -When `backend-tornado` is on the application classpath, `PureJavaBackend.loadAutomatic(...)` -discovers it through Java's standard `ServiceLoader`. ModelJars uses that automatic -entry point for artifacts qualified for the Java backend, so adding the optional -dependency and launching with TornadoVM is sufficient; application inference code -does not change. Without the optional module or a qualified device, the same call -loads the Vector API backend. +`backend-cuda` optionally executes the K-quant projections and grouped-query +decode attention on a qualified NVIDIA GPU, using Models-owned Rust kernels +compiled to PTX. The model parser, tokenizer, transformer graph, KV cache, +sampling and generation loop remain inside Models. There is no external +inference server, no vendor math library and no handwritten CUDA C++. -== Open And Inspect The Backend +Models ships *one* GPU implementation. A TornadoVM-based arm existed until +0.3.53 and was removed: on the same RTX 4090 and the same model it reached +6.49 tok/s decode over 36 failed `cuLaunchKernel` calls, against 31.26 tok/s +for this path, and it required a separately installed device runtime that a +Maven artifact cannot carry. The measurement is retained in +`benchmark-results/2026-10-07-tornado-removal`. -[source,java] ----- -var options = TornadoBackendOptions.defaults(); +== Nothing activates it by accident -try (var runtime = TornadoBackend.open(modelPath, backendConfiguration, options)) { - var backend = runtime.backend(); - var status = runtime.status(); - - System.out.printf("device=%s accelerated=%s readiness=%s%n", - status.device(), status.accelerated(), status.readinessTime()); - // Pass backend to GenerationLoop or the normal Models runtime pipeline. -} ----- - -Applications that do not need the detailed readiness status can use the -service-loaded entry point directly: +The jar carries no `META-INF/services` entry, so no `ServiceLoader` discovers +it. An application opens the accelerator explicitly: [source,java] ---- -try (var backend = PureJavaBackend.loadAutomatic(modelPath, backendConfiguration)) { - // Use the normal Models generation pipeline. -} +CudaGgufBatchedMatrixKernel.Status status = CudaGgufBatchedMatrixKernel.open(); ---- -Defaults perform eager readiness, accelerate eligible prefill and decode -projections, and safely load `PureJavaBackend` when acceleration is not -available. A qualification or deployment gate can set `requireAccelerator` to -`true`; in that mode an ineligible device or initialization failure stops the -load instead of falling back. +`open()` never throws for an absent, old or ineligible device; it returns a +status explaining why. `-Dmodels.cuda.disabled=true` refuses unconditionally, +and `-Dmodels.cuda.attention.disabled=true` ablates only the attention kernel, +counting the refusal rather than hiding it. -Readiness is intentional. TornadoVM retains compiled code with each execution -plan, so Models constructs and compiles the fixed shapes before accepting a -visible request. `TornadoBackendStatus` reports that cost and the selected -device. +== One module for every qualifying device -A mixed-format model can look accelerated while one of its formats silently -falls back, so `TornadoBackendRuntime` also reports how many projections reached -the device per GGUF weight format. A grouped dispatch counts once per matrix. +PTX is device code, so a single `sm_80` module serves every device of compute +capability 8.0 or above. It travels inside the jar under +`META-INF/models/cuda/` with a SHA-256 the Java loader recomputes on load, and +an ABI field that must match the Java binding. No CUDA toolkit is needed to +build or to run; only a shipping `libcuda.so.1`. -[source,java] ----- -try (var runtime = TornadoBackend.open(modelPath, backendConfiguration, options)) { - // Run a prefill first; counts are empty until projections are dispatched. - System.out.println(runtime.routedProjectionsByFormat()); // {Q4_K=1344, Q6_K=64} - System.out.println(runtime.projectionPlanCount()); -} ----- +== Measured gates -== Accelerator Profile Gate - -`models-bench` runs prefill and decode for a GGUF file on the Tornado backend and -writes the device, readiness cost, throughput, and per-format routing counts as -JSON: +Run them with `models-bench cuda-kernel-gate`, which has three modes and one +report shape: [source,shell] ---- -./gradlew :models-bench:installDist +models-bench cuda-kernel-gate --mode capability --require-device true \ + --report capability.json --models-revision "$(git rev-parse HEAD)" -lib=models-bench/build/install/models-bench/lib -classpath=$(printf '%s:' "$lib"/*.jar) +models-bench cuda-kernel-gate --mode parity --model model.gguf \ + --prompts prompts.txt --max-tokens 64 --require-device true \ + --report parity.json --models-revision "$(git rev-parse HEAD)" -tornado -cp "$classpath" \ - --params="accelerator-profile --model /path/to/model-Q4_K_M.gguf \ - --tokens 64 --batch 32 --require true \ - --output build/reports/inference/accelerator-profile.json" \ - com.integrallis.models.bench.InferenceBenchmarkCli +models-bench cuda-kernel-gate --mode decode --model model.gguf \ + --prompts prompts.txt --max-tokens 64 --warmup-tokens 16 --arm both \ + --require-device true --report decode.json \ + --models-revision "$(git rev-parse HEAD)" ---- -`--require true` makes an unavailable or ineligible accelerator a hard failure -rather than a silent Vector API fallback. Without a qualified device the command -still runs and the report records `accelerated=false` with the reason. - -== Measured Gates +`--arm both` measures the accelerated arm and the Vector API control in the +same process on the same host, which is the only denominator that counts. +`--models-revision` requires a full 40-character SHA. -The release candidate preserved the exact greedy CPU token sequence on both -qualified profiles using Qwen3 0.6B Q4_0. These are Q4_0 numbers; the K-quant -kernels have not yet been measured on hardware, and no throughput claim is made -for them here. +Measured on an RTX 4090 (compute capability 8.9), Granite 4.1 3B Q4_K_M, 20 +prompts: -[cols="1,1,1,1",options="header"] +[cols="2,1,3"] |=== -|Profile |Warm prefill |Median decode |Eager readiness -|NVIDIA A16-2Q |4.938 s (`1.93x` CPU) |73.991 ms (`1.21x` CPU) |14.278 s -|NVIDIA A40-4Q |2.150 s (`4.40x` CPU) |65.767 ms (`1.32x` CPU) |13.554 s +|Gate |Result |Evidence + +|G1 exact token parity +|*passed* +|1280 of 1280 token ids identical, `benchmark-results/2026-10-07-g1-parity` + +|G4 decode speed +|*passed* +|6.184x against a 3.00x threshold, 31.26 against 5.06 tok/s, + `benchmark-results/2026-10-07-g4-dualpath` + +|G2 routing observability +|passed +|2,405,576 accelerated operations, zero declined projections + +|G5 startup honesty +|passed +|344 ms readiness against a 120,000 ms ceiling |=== -These are qualification measurements for the tested virtual-GPU profiles, not -universal performance promises. Application latency depends on device shape, -CPU allocation, prompt length, model size, and driver/runtime versions. +G1 carries no tolerance. Each qualifying hardware profile and each +architecture family earns its own run before the result generalises: the figures +above are one device and one model. + +== What routes, and how to tell + +The report records accelerated operations per weight format and stage, plus +`declinedProjections` keyed by format, stage, reason and shape. That second map +exists because an unimplemented path is not a refusal: the FFN gate and up +projections ran on the CPU through an entire measurement because +`multiplyDual` was missing, 53.3% of a layer's projection arithmetic, and an +empty refusals map read as "everything ran on the device". A projection that +answers no to `isEligible` now increments a counter that names it. + +Read `launchesPerDecodeStep`, `transfersPerDecodeStep` and +`activationBytesPerDecodeStep` together with the throughput. The decode path +issues roughly 321 launches and 522 transfers per token on a 40-block model, +because every projection result returns to a Java array; that is the standing +overhead and the next structural lever against it is device-resident +activations. diff --git a/docs/content/modules/ROOT/pages/index.adoc b/docs/content/modules/ROOT/pages/index.adoc index 36f76c4f..12f5e995 100644 --- a/docs/content/modules/ROOT/pages/index.adoc +++ b/docs/content/modules/ROOT/pages/index.adoc @@ -27,10 +27,11 @@ the same Java pipeline but replaces selected, measured bottleneck kernels with a small Models-owned Rust library through Java's Foreign Function and Memory (FFM) API. It does not embed or invoke llama.cpp or Ollama. -The optional `backend-tornado` module keeps that graph and its projection -kernels in Java while TornadoVM compiles qualified Q4_0 work for an NVIDIA GPU. -It performs an explicit readiness pass and falls back to the Vector API when -the device, memory capacity, or artifact is not eligible. +The optional `backend-cuda` module keeps that graph in Java while Models-owned +Rust kernels, compiled to PTX and carried inside the jar, run the K-quant +projections and grouped-query decode attention on an NVIDIA GPU. It performs an +explicit readiness pass and falls back to the Vector API when the device, memory +capacity, or artifact is not eligible. New to local model execution? Start with xref:concepts.adoc[Inference Concepts]. @@ -104,8 +105,9 @@ image::models-0001.png[Models runtime architecture,align=center] |`backend-java` |GGUF, supported Safetensors, and CACT model loading with Java 25 Vector API execution -|`backend-tornado` -|Optional Java-authored Q4_0 projection acceleration on qualified NVIDIA GPUs +|`backend-cuda` +|Optional Rust-authored PTX projection and decode-attention acceleration on +qualified NVIDIA GPUs |`backend-native` |The Java backend with selective Rust/FFM bottleneck-kernel substitution diff --git a/docs/content/modules/ROOT/pages/modules.adoc b/docs/content/modules/ROOT/pages/modules.adoc index 331b1433..b548e27d 100644 --- a/docs/content/modules/ROOT/pages/modules.adoc +++ b/docs/content/modules/ROOT/pages/modules.adoc @@ -25,8 +25,9 @@ execution across local and hosted clients |`backend-java` |GGUF parser, tokenizers, model graph, KV cache, and Java kernels -|`backend-tornado` |Optional Java-authored Q4_0 projection kernels compiled -for qualified NVIDIA GPUs by TornadoVM, with eager readiness and CPU fallback +|`backend-cuda` |Optional Models-owned Rust kernels compiled to PTX for +qualified NVIDIA GPUs: K-quant projections and grouped-query decode attention, +opt-in, with CPU fallback |`backend-native` |Java 25 and Vector API backend with selective Rust kernel acceleration through FFM @@ -77,7 +78,7 @@ models-api <- models-runtime <- models <- backend-java <- vectors-core - <- backend-tornado <- TornadoVM device compiler/runtime + <- backend-cuda <- Models Rust PTX kernels through FFM <- backend-native <- Models Rust kernels through FFM <- backend-apple <- Apple's Foundation Models framework through FFM <- models-router <- vectors-db @@ -105,7 +106,7 @@ Application dependencies use: * `com.integrallis:models` * `com.integrallis:backend-java` -* `com.integrallis:backend-tornado` +* `com.integrallis:backend-cuda` * `com.integrallis:backend-native` * `com.integrallis:backend-apple` * `com.integrallis:models-audio` diff --git a/models-accelerator-bench/README.md b/models-accelerator-bench/README.md index 12a8c764..0bb77289 100644 --- a/models-accelerator-bench/README.md +++ b/models-accelerator-bench/README.md @@ -1,9 +1,10 @@ # Java accelerator experiments This private Gradle module retains accelerator experiments, rejected candidates, and release gates. -The qualified Q4 projection provider now lives in the optional published `backend-tornado` module. -Its kernels are Java source compiled for an accelerator by TornadoVM; there is no external -inference server and no handwritten CUDA or Rust shim. +The GPU provider is the optional published `backend-cuda` module: Models-owned Rust kernels +compiled to PTX, carried inside the jar, with no external inference server and no handwritten CUDA +C++. The TornadoVM arm this module used to benchmark was removed in 0.3.53 -- see +`benchmark-results/2026-10-07-tornado-removal`. What remains here is the KV-ridge experiment. The first experiment targets quantized matrix projection because profiling and the existing Models execution seams identify it as the dominant reusable operation. It currently contains: diff --git a/models-accelerator-bench/build.gradle.kts b/models-accelerator-bench/build.gradle.kts index 7ab9a062..cf06e014 100644 --- a/models-accelerator-bench/build.gradle.kts +++ b/models-accelerator-bench/build.gradle.kts @@ -43,11 +43,8 @@ application { dependencies { implementation(project(":backend-java")) - implementation(project(":backend-tornado")) implementation(project(":models-runtime")) implementation("com.integrallis:vectors-core:${providers.gradleProperty("vectorsVersion").get()}") - implementation("io.github.beehive-lab:tornado-api:5.2.0-jdk25") - implementation("io.github.beehive-lab:tornado-runtime:5.2.0-jdk25") testImplementation("org.junit.jupiter:junit-jupiter:5.11.4") testImplementation("org.assertj:assertj-core:3.27.2") testRuntimeOnly("org.junit.platform:junit-platform-launcher:1.11.4") diff --git a/models-accelerator-bench/src/main/java/com/integrallis/models/accelerator/CausalAttentionExperiment.java b/models-accelerator-bench/src/main/java/com/integrallis/models/accelerator/CausalAttentionExperiment.java deleted file mode 100644 index dcdb528e..00000000 --- a/models-accelerator-bench/src/main/java/com/integrallis/models/accelerator/CausalAttentionExperiment.java +++ /dev/null @@ -1,198 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.accelerator; - -import java.util.Arrays; -import java.util.Locale; -import java.util.SplittableRandom; - -/** Device correctness and latency gate for Qwen3-shaped causal grouped-query attention. */ -public final class CausalAttentionExperiment { - private static final int EXECUTION_BATCH = 32; - private static final int HEADS = 16; - private static final int KV_HEADS = 8; - private static final int HEAD_LENGTH = 64; - private static final int QUERY_DIM = HEADS * HEAD_LENGTH; - private static final int KV_DIM = KV_HEADS * HEAD_LENGTH; - private static final int MAX_SEQUENCE = 512; - private static final int CHUNKS = 8; - private static final int MEASUREMENTS = 7; - private static final double MAX_RELATIVE_L2 = 2.0e-5; - - private CausalAttentionExperiment() {} - - public static void main(String[] args) { - if (args.length != 0) { - throw new IllegalArgumentException("usage: CausalAttentionExperiment"); - } - Chunk[] chunks = chunks(); - float[][] expected = referenceOutputs(chunks); - float[][] actual = new float[CHUNKS][EXECUTION_BATCH * QUERY_DIM]; - - try (TornadoCausalAttentionPlan plan = - new TornadoCausalAttentionPlan( - "qwen3-causal-attention", - EXECUTION_BATCH, - HEADS, - KV_HEADS, - HEAD_LENGTH, - HEAD_LENGTH, - MAX_SEQUENCE, - 0)) { - long coldStarted = System.nanoTime(); - executeSequence(plan, chunks, actual); - double coldMillis = elapsedMillis(coldStarted); - double maximumError = 0.0; - for (int chunk = 0; chunk < CHUNKS; chunk++) { - maximumError = - Math.max( - maximumError, Q4ProjectionExperiment.relativeL2(expected[chunk], actual[chunk])); - } - if (maximumError > MAX_RELATIVE_L2) { - throw new IllegalStateException( - "attention result failed relative-L2 gate: " + maximumError); - } - - double[] gpuMillis = new double[MEASUREMENTS]; - for (int measurement = 0; measurement < MEASUREMENTS; measurement++) { - long started = System.nanoTime(); - executeSequence(plan, chunks, actual); - gpuMillis[measurement] = elapsedMillis(started); - } - double[] cpuMillis = new double[MEASUREMENTS]; - for (int measurement = 0; measurement < MEASUREMENTS; measurement++) { - long started = System.nanoTime(); - referenceOutputs(chunks); - cpuMillis[measurement] = elapsedMillis(started); - } - System.out.printf( - Locale.ROOT, - "Qwen3 attention chunks=%d tokens=%d relativeL2=%.8g cold=%.3f ms GPU-p50=%.3f ms CPU-p50=%.3f ms speedup=%.2fx calls=%d%n", - CHUNKS, - CHUNKS * EXECUTION_BATCH, - maximumError, - coldMillis, - median(gpuMillis), - median(cpuMillis), - median(cpuMillis) / median(gpuMillis), - plan.calls()); - } - } - - private static void executeSequence( - TornadoCausalAttentionPlan plan, Chunk[] chunks, float[][] outputs) { - for (int chunk = 0; chunk < chunks.length; chunk++) { - Chunk inputs = chunks[chunk]; - plan.execute( - inputs.query(), - inputs.key(), - inputs.value(), - outputs[chunk], - EXECUTION_BATCH, - chunk * EXECUTION_BATCH); - } - } - - private static Chunk[] chunks() { - Chunk[] chunks = new Chunk[CHUNKS]; - for (int chunk = 0; chunk < chunks.length; chunk++) { - long seed = 101L + chunk * 10L; - chunks[chunk] = - new Chunk( - randomFloats(EXECUTION_BATCH * QUERY_DIM, seed), - randomFloats(EXECUTION_BATCH * KV_DIM, seed + 1), - randomFloats(EXECUTION_BATCH * KV_DIM, seed + 2)); - } - return chunks; - } - - private static float[] randomFloats(int length, long seed) { - SplittableRandom random = new SplittableRandom(seed); - float[] values = new float[length]; - for (int index = 0; index < values.length; index++) { - values[index] = (float) random.nextDouble(-1.0, 1.0); - } - return values; - } - - private static float[][] referenceOutputs(Chunk[] chunks) { - float[] keyCache = new float[MAX_SEQUENCE * KV_DIM]; - float[] valueCache = new float[MAX_SEQUENCE * KV_DIM]; - float[][] outputs = new float[chunks.length][EXECUTION_BATCH * QUERY_DIM]; - for (int chunk = 0; chunk < chunks.length; chunk++) { - int startPosition = chunk * EXECUTION_BATCH; - Chunk inputs = chunks[chunk]; - for (int batch = 0; batch < EXECUTION_BATCH; batch++) { - System.arraycopy( - inputs.key(), batch * KV_DIM, keyCache, (startPosition + batch) * KV_DIM, KV_DIM); - System.arraycopy( - inputs.value(), batch * KV_DIM, valueCache, (startPosition + batch) * KV_DIM, KV_DIM); - } - referenceChunk(inputs.query(), keyCache, valueCache, outputs[chunk], startPosition); - } - return outputs; - } - - private static void referenceChunk( - float[] query, float[] keyCache, float[] valueCache, float[] output, int startPosition) { - float scale = (float) (1.0 / Math.sqrt(HEAD_LENGTH)); - int groupSize = HEADS / KV_HEADS; - float[] scores = new float[MAX_SEQUENCE]; - for (int batch = 0; batch < EXECUTION_BATCH; batch++) { - int position = startPosition + batch; - for (int head = 0; head < HEADS; head++) { - int kvHead = head / groupSize; - float maximum = Float.NEGATIVE_INFINITY; - for (int cached = 0; cached <= position; cached++) { - float dot = 0.0f; - for (int column = 0; column < HEAD_LENGTH; column++) { - dot += - query[batch * QUERY_DIM + head * HEAD_LENGTH + column] - * keyCache[cached * KV_DIM + kvHead * HEAD_LENGTH + column]; - } - scores[cached] = dot * scale; - maximum = Math.max(maximum, scores[cached]); - } - float denominator = 0.0f; - for (int cached = 0; cached <= position; cached++) { - scores[cached] = (float) Math.exp(scores[cached] - maximum); - denominator += scores[cached]; - } - int outputOffset = batch * QUERY_DIM + head * HEAD_LENGTH; - for (int column = 0; column < HEAD_LENGTH; column++) { - float weighted = 0.0f; - for (int cached = 0; cached <= position; cached++) { - weighted += - scores[cached] * valueCache[cached * KV_DIM + kvHead * HEAD_LENGTH + column]; - } - output[outputOffset + column] = weighted / denominator; - } - } - } - } - - private static double median(double[] values) { - double[] sorted = values.clone(); - Arrays.sort(sorted); - return sorted[sorted.length / 2]; - } - - private static double elapsedMillis(long started) { - return (System.nanoTime() - started) / 1_000_000.0; - } - - private record Chunk(float[] query, float[] key, float[] value) {} -} diff --git a/models-accelerator-bench/src/main/java/com/integrallis/models/accelerator/CausalAttentionKernel.java b/models-accelerator-bench/src/main/java/com/integrallis/models/accelerator/CausalAttentionKernel.java deleted file mode 100644 index 9cfec614..00000000 --- a/models-accelerator-bench/src/main/java/com/integrallis/models/accelerator/CausalAttentionKernel.java +++ /dev/null @@ -1,238 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.accelerator; - -import uk.ac.manchester.tornado.api.KernelContext; -import uk.ac.manchester.tornado.api.annotations.Parallel; -import uk.ac.manchester.tornado.api.math.TornadoMath; -import uk.ac.manchester.tornado.api.types.arrays.FloatArray; -import uk.ac.manchester.tornado.api.types.arrays.IntArray; - -/** Java-authored fixed-shape grouped-query attention experiment with a device-resident KV cache. */ -final class CausalAttentionKernel { - private static final int START_POSITION_INDEX = 0; - private static final int ACTUAL_BATCH_INDEX = 1; - - private CausalAttentionKernel() {} - - /** Stores the current prompt chunk in a persistent, position-major KV cache. */ - static void store( - IntArray state, - FloatArray key, - FloatArray value, - FloatArray keyCache, - FloatArray valueCache, - int keyDim, - int valueDim) { - int executionBatchSize = key.getSize() / keyDim; - int entriesPerBatch = keyDim + valueDim; - for (@Parallel int entry = 0; entry < executionBatchSize * entriesPerBatch; entry++) { - int batch = entry / entriesPerBatch; - int component = entry - batch * entriesPerBatch; - if (batch < state.get(ACTUAL_BATCH_INDEX)) { - int position = state.get(START_POSITION_INDEX) + batch; - if (component < keyDim) { - keyCache.set(position * keyDim + component, key.get(batch * keyDim + component)); - } else { - int valueComponent = component - keyDim; - valueCache.set( - position * valueDim + valueComponent, value.get(batch * valueDim + valueComponent)); - } - } - } - } - - /** Computes causal grouped-query attention with one independent work item per token/head pair. */ - static void attend( - IntArray state, - FloatArray query, - FloatArray keyCache, - FloatArray valueCache, - FloatArray scores, - FloatArray output, - int numHeads, - int numKvHeads, - int keyLength, - int valueLength, - int keyDim, - int valueDim, - int maxSequenceLength, - int slidingWindow) { - int queryDim = numHeads * keyLength; - int outputDim = numHeads * valueLength; - int executionBatchSize = query.getSize() / queryDim; - int groupSize = numHeads / numKvHeads; - float scale = 1.0f / TornadoMath.sqrt(keyLength); - for (@Parallel int task = 0; task < executionBatchSize * numHeads; task++) { - int batch = task / numHeads; - int head = task - batch * numHeads; - int outputOffset = batch * outputDim + head * valueLength; - if (batch < state.get(ACTUAL_BATCH_INDEX)) { - int position = state.get(START_POSITION_INDEX) + batch; - int firstPosition = slidingWindow > 0 ? Math.max(0, position - slidingWindow + 1) : 0; - int kvHead = head / groupSize; - int queryOffset = batch * queryDim + head * keyLength; - int scoreOffset = task * maxSequenceLength; - float maximum = Float.NEGATIVE_INFINITY; - for (int cachedPosition = firstPosition; cachedPosition <= position; cachedPosition++) { - int keyOffset = cachedPosition * keyDim + kvHead * keyLength; - float score = 0.0f; - for (int column = 0; column < keyLength; column++) { - score += query.get(queryOffset + column) * keyCache.get(keyOffset + column); - } - score *= scale; - scores.set(scoreOffset + cachedPosition, score); - maximum = Math.max(maximum, score); - } - - float denominator = 0.0f; - for (int cachedPosition = firstPosition; cachedPosition <= position; cachedPosition++) { - float probability = TornadoMath.exp(scores.get(scoreOffset + cachedPosition) - maximum); - scores.set(scoreOffset + cachedPosition, probability); - denominator += probability; - } - - for (int column = 0; column < valueLength; column++) { - float weightedValue = 0.0f; - for (int cachedPosition = firstPosition; cachedPosition <= position; cachedPosition++) { - int valueOffset = cachedPosition * valueDim + kvHead * valueLength; - weightedValue += - scores.get(scoreOffset + cachedPosition) * valueCache.get(valueOffset + column); - } - output.set(outputOffset + column, weightedValue / denominator); - } - } else { - for (int column = 0; column < valueLength; column++) { - output.set(outputOffset + column, 0.0f); - } - } - } - } - - /** - * Workgroup-tiled attention adapted from the local GPULlama3 reference implementation. - * - *

One workgroup owns a token/head pair. K/V rows are staged through local memory and softmax - * is accumulated online across 16-position tiles. - */ - static void attendTiled( - KernelContext context, - IntArray state, - FloatArray query, - FloatArray keyCache, - FloatArray valueCache, - FloatArray output, - int numHeads, - int keyLength, - int keyDim, - int groupSize, - int queryDim) { - int localThread = context.localIdx; - int localSize = context.localGroupSizeX; - int task = context.groupIdx; - int batch = task / numHeads; - int head = task - batch * numHeads; - int outputOffset = batch * queryDim + head * keyLength; - int tileSize = 16; - - float[] sharedQuery = context.allocateFloatLocalArray(keyLength); - float[] keyTile = context.allocateFloatLocalArray(tileSize * keyLength); - float[] valueTile = context.allocateFloatLocalArray(tileSize * keyLength); - float[] scoreTile = context.allocateFloatLocalArray(tileSize); - float[] maximumHolder = context.allocateFloatLocalArray(1); - - if (batch < state.get(ACTUAL_BATCH_INDEX)) { - int position = state.get(START_POSITION_INDEX) + batch; - int kvHead = head / groupSize; - int queryOffset = batch * queryDim + head * keyLength; - for (int column = localThread; column < keyLength; column += localSize) { - sharedQuery[column] = query.get(queryOffset + column); - } - context.localBarrier(); - - float maximum = Float.NEGATIVE_INFINITY; - float denominator = 0.0f; - float[] accumulator = new float[keyLength]; - for (int column = 0; column < keyLength; column++) { - accumulator[column] = 0.0f; - } - - for (int tileStart = 0; tileStart <= position; tileStart += tileSize) { - int tileEnd = Math.min(tileStart + tileSize - 1, position); - for (int cachedPosition = tileStart + localThread; - cachedPosition <= tileEnd; - cachedPosition += localSize) { - int tilePosition = cachedPosition - tileStart; - int tileOffset = tilePosition * keyLength; - int cacheOffset = cachedPosition * keyDim + kvHead * keyLength; - for (int column = 0; column < keyLength; column++) { - keyTile[tileOffset + column] = keyCache.get(cacheOffset + column); - valueTile[tileOffset + column] = valueCache.get(cacheOffset + column); - } - } - context.localBarrier(); - - for (int cachedPosition = tileStart + localThread; - cachedPosition <= tileEnd; - cachedPosition += localSize) { - int tilePosition = cachedPosition - tileStart; - float score = 0.0f; - for (int column = 0; column < keyLength; column++) { - score += sharedQuery[column] * keyTile[tilePosition * keyLength + column]; - } - scoreTile[tilePosition] = score / TornadoMath.sqrt(keyLength); - } - context.localBarrier(); - - float tileMaximum = Float.NEGATIVE_INFINITY; - for (int tilePosition = 0; tilePosition <= tileEnd - tileStart; tilePosition++) { - tileMaximum = Math.max(tileMaximum, scoreTile[tilePosition]); - } - if (localThread == 0) { - maximumHolder[0] = tileMaximum; - } - context.localBarrier(); - - float newMaximum = Math.max(maximum, maximumHolder[0]); - if (maximum != Float.NEGATIVE_INFINITY) { - float previousScale = TornadoMath.exp(maximum - newMaximum); - denominator *= previousScale; - for (int column = 0; column < keyLength; column++) { - accumulator[column] *= previousScale; - } - } - maximum = newMaximum; - for (int tilePosition = 0; tilePosition <= tileEnd - tileStart; tilePosition++) { - float probability = TornadoMath.exp(scoreTile[tilePosition] - maximum); - denominator += probability; - for (int column = 0; column < keyLength; column++) { - accumulator[column] += probability * valueTile[tilePosition * keyLength + column]; - } - } - context.localBarrier(); - } - - float inverseDenominator = denominator > 0.0f ? 1.0f / denominator : 0.0f; - for (int column = localThread; column < keyLength; column += localSize) { - output.set(outputOffset + column, accumulator[column] * inverseDenominator); - } - } else { - for (int column = localThread; column < keyLength; column += localSize) { - output.set(outputOffset + column, 0.0f); - } - } - } -} diff --git a/models-accelerator-bench/src/main/java/com/integrallis/models/accelerator/DeviceInventoryExperiment.java b/models-accelerator-bench/src/main/java/com/integrallis/models/accelerator/DeviceInventoryExperiment.java deleted file mode 100644 index 279ab806..00000000 --- a/models-accelerator-bench/src/main/java/com/integrallis/models/accelerator/DeviceInventoryExperiment.java +++ /dev/null @@ -1,68 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.accelerator; - -import com.integrallis.models.backend.tornado.AcceleratorEligibility; -import com.integrallis.models.backend.tornado.TornadoRuntimeDevices; -import java.io.IOException; -import java.io.UncheckedIOException; -import java.nio.file.Files; -import java.nio.file.Path; -import java.util.Arrays; -import java.util.List; -import java.util.Locale; - -/** Prints the exact runtime facts used by automatic accelerator eligibility and fallback. */ -public final class DeviceInventoryExperiment { - private static final double MIB = 1024.0 * 1024.0; - - private DeviceInventoryExperiment() {} - - public static void main(String[] args) { - if (args.length < 1 || args.length > 2 || (args.length == 2 && !"--decode".equals(args[1]))) { - throw new IllegalArgumentException( - "usage: DeviceInventoryExperiment [--decode]"); - } - Path model = Path.of(args[0]).toAbsolutePath().normalize(); - long modelBytes; - try { - modelBytes = Files.size(model); - } catch (IOException exception) { - throw new UncheckedIOException("could not read model size: " + model, exception); - } - boolean decode = Arrays.asList(args).contains("--decode"); - List devices = TornadoRuntimeDevices.discover(); - for (AcceleratorEligibility.DeviceCapabilities device : devices) { - System.out.printf( - Locale.ROOT, - "device=%s backend=%s type=%s memory=%.1f MiB maxAllocation=%.1f MiB%n", - device.name(), - device.backend(), - device.type(), - device.globalMemoryBytes() / MIB, - device.maxAllocationBytes() / MIB); - } - AcceleratorEligibility.Decision decision = - AcceleratorEligibility.select(devices, modelBytes, decode); - System.out.printf( - Locale.ROOT, - "eligible=%s selected=%s required=%.1f MiB reason=%s%n", - decision.eligible(), - decision.device() == null ? "none" : decision.device().name(), - decision.requiredBytes() / MIB, - decision.reason()); - } -} diff --git a/models-accelerator-bench/src/main/java/com/integrallis/models/accelerator/LargeModelBudgetExperiment.java b/models-accelerator-bench/src/main/java/com/integrallis/models/accelerator/LargeModelBudgetExperiment.java deleted file mode 100644 index f6660d47..00000000 --- a/models-accelerator-bench/src/main/java/com/integrallis/models/accelerator/LargeModelBudgetExperiment.java +++ /dev/null @@ -1,212 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.accelerator; - -import com.integrallis.models.backend.tornado.AcceleratorEligibility; -import com.integrallis.models.backend.tornado.DeviceBudget; -import com.integrallis.models.backend.tornado.DeviceMemoryRequest; -import com.integrallis.models.backend.tornado.PlanShapeStrategy; -import com.integrallis.models.backend.tornado.TornadoRuntimeDevices; -import java.util.ArrayList; -import java.util.LinkedHashMap; -import java.util.List; -import java.util.Locale; -import java.util.Map; - -/** - * Prints the accelerator device budget for a model shape against a table of devices. - * - *

Runs without a GPU: the device table is data, so a new question about a new device or context - * length never needs a recompile. When a TornadoVM runtime is present its discovered devices are - * appended to the table. - * - *

Example: - * - *

{@code
- * LargeModelBudgetExperiment --preset gemma4-26b-a4b --context 4096,32768,262144 \
- *     --device "A40 24 GB:23:8" --device "L40S 48 GB:44:8" --device "H100 80 GB:79:32"
- * }
- */ -public final class LargeModelBudgetExperiment { - - private static final long MIB = 1024L * 1024L; - private static final long GIB = 1024L * MIB; - - private LargeModelBudgetExperiment() {} - - public static void main(String[] args) { - Map options = parse(args); - Shape shape = shape(options); - List contexts = contexts(options); - List devices = devices(options); - - System.out.printf( - Locale.ROOT, - "model=%s weights=%s shapes=%d plans=%d planScratch=%s largestAllocation=%s%n", - shape.label(), - DeviceBudget.gib(shape.weightBytes()), - shape.retainedShapes(), - shape.retainedPlanCount(), - DeviceBudget.gib(shape.planScratchBytes()), - DeviceBudget.gib(shape.largestAllocationBytes())); - System.out.println(); - - for (int context : contexts) { - DeviceMemoryRequest request = shape.request(context); - DeviceBudget shipped = - AcceleratorEligibility.budget(request, PlanShapeStrategy.PER_SHAPE_WHOLE_MODEL); - DeviceBudget shared = - AcceleratorEligibility.budget(request, PlanShapeStrategy.SHARED_WEIGHT_UPLOAD); - System.out.printf( - Locale.ROOT, - "context=%d kv=%s shipped=%s shared=%s hostCopies=%s readiness=%d s%n", - context, - DeviceBudget.gib(request.deviceKvCacheBytes()), - DeviceBudget.gib(shipped.totalBytes()), - DeviceBudget.gib(shared.totalBytes()), - DeviceBudget.gib(shipped.hostWeightCopyBytes()), - shipped.estimatedPlanCompileTime().toSeconds()); - for (AcceleratorEligibility.DeviceCapabilities device : devices) { - AcceleratorEligibility.Decision decision = - AcceleratorEligibility.select(List.of(device), request); - System.out.printf( - Locale.ROOT, - " %-16s global=%s eligible=%s %s%n", - device.name(), - DeviceBudget.gib(device.globalMemoryBytes()), - decision.eligible(), - decision.eligible() ? "" : decision.reason()); - } - System.out.println(); - } - } - - private record Shape( - String label, - long weightBytes, - int retainedShapes, - int retainedPlanCount, - long planScratchBytes, - long largestAllocationBytes, - KvGeometry kv) { - - private DeviceMemoryRequest request(int contextLength) { - return DeviceMemoryRequest.detailed(label) - .weightBytes(weightBytes) - .retainedShapes(retainedShapes) - .retainedPlanCount(retainedPlanCount) - .planScratchBytes(planScratchBytes) - .largestAllocationBytes(largestAllocationBytes) - .deviceKvCacheBytes(kv.bytes(contextLength)) - .build(); - } - } - - /** Ring-bounded plus context-linear KV, matching {@code LayeredKvCache}. */ - private record KvGeometry(long ringBytes, long bytesPerToken, int maxContext) { - private long bytes(int contextLength) { - return ringBytes + bytesPerToken * Math.min(contextLength, maxContext); - } - } - - private static Shape shape(Map options) { - String preset = options.getOrDefault("preset", "gemma4-26b-a4b"); - Shape base = - switch (preset) { - // Gemma 4 26B-A4B IT Q4_K_M, every constant read from the Models repository: - // resident and routed-expert bytes from Gemma4LargeModelFixtureSlowTest; 30 layers with - // 25 sliding (keyDim 2048, ring 1024+1) and 5 full (keyDim 1024) from Gemma4ConfigTest - // and Gemma4KvCache; plans = 30 x (128 experts x 2 + 4 resident) x 2 retained shapes. - case "gemma4-26b-a4b" -> - new Shape( - "gemma-4-26B-A4B-it-Q4_K_M.gguf", - 1_650_027_640L + 15_130_165_248L, - 2, - 30 * (128 * 2 + 4) * 2, - 2_733_284_400L, - 262_144L * 2_816L / 32L * 34L, - new KvGeometry(25L * 2 * 2_048 * 4 * 1_025, 5L * 2 * 1_024 * 4, 262_144)); - // Qwen3 0.6B Q4_0: the published A16-2Q / A40-4Q continuity control. - case "qwen3-0.6b" -> - new Shape( - "Qwen3-0.6B-Q4_0.gguf", - 428_970_080L, - 2, - 223, - 0L, - 0L, - new KvGeometry(0L, 0L, 40_960)); - default -> throw new IllegalArgumentException("unknown preset: " + preset); - }; - return new Shape( - options.getOrDefault("label", base.label()), - longOption(options, "weight-bytes", base.weightBytes()), - (int) longOption(options, "shapes", base.retainedShapes()), - (int) longOption(options, "plans", base.retainedPlanCount()), - longOption(options, "plan-scratch", base.planScratchBytes()), - longOption(options, "largest-allocation", base.largestAllocationBytes()), - base.kv()); - } - - private static List contexts(Map options) { - String raw = options.getOrDefault("context", "4096,32768,131072,262144"); - List contexts = new ArrayList<>(); - for (String token : raw.split(",", -1)) { - contexts.add(Integer.parseInt(token.strip())); - } - return contexts; - } - - private static List devices( - Map options) { - List devices = new ArrayList<>(); - String raw = options.getOrDefault("device", "A40 24 GB:23:8|L40S 48 GB:44:8|H100 80 GB:79:32"); - for (String entry : raw.split("\\|", -1)) { - String[] parts = entry.split(":", -1); - if (parts.length != 3) { - throw new IllegalArgumentException("device must be name:globalGiB:maxAllocGiB: " + entry); - } - devices.add( - new AcceleratorEligibility.DeviceCapabilities( - parts[0].strip(), - "PTX", - "GPU", - Long.parseLong(parts[1].strip()) * GIB, - Long.parseLong(parts[2].strip()) * GIB)); - } - devices.addAll(TornadoRuntimeDevices.discover()); - return devices; - } - - private static long longOption(Map options, String name, long fallback) { - String value = options.get(name); - return value == null ? fallback : Long.parseLong(value.strip()); - } - - private static Map parse(String[] args) { - Map options = new LinkedHashMap<>(); - for (int index = 0; index < args.length; index++) { - String argument = args[index]; - if (!argument.startsWith("--") || index + 1 >= args.length) { - throw new IllegalArgumentException("expected --option value pairs, found: " + argument); - } - String key = argument.substring(2); - String value = args[++index]; - options.merge(key, value, (existing, added) -> existing + "|" + added); - } - return options; - } -} diff --git a/models-accelerator-bench/src/main/java/com/integrallis/models/accelerator/Q4GroupedProjectionExperiment.java b/models-accelerator-bench/src/main/java/com/integrallis/models/accelerator/Q4GroupedProjectionExperiment.java deleted file mode 100644 index 61e6f707..00000000 --- a/models-accelerator-bench/src/main/java/com/integrallis/models/accelerator/Q4GroupedProjectionExperiment.java +++ /dev/null @@ -1,195 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.accelerator; - -import com.integrallis.models.backend.tornado.Q4ProjectionKernel; -import java.util.Locale; -import uk.ac.manchester.tornado.api.TaskGraph; -import uk.ac.manchester.tornado.api.TornadoExecutionPlan; -import uk.ac.manchester.tornado.api.enums.DataTransferMode; -import uk.ac.manchester.tornado.api.exceptions.TornadoExecutionPlanException; -import uk.ac.manchester.tornado.api.types.arrays.ByteArray; -import uk.ac.manchester.tornado.api.types.arrays.FloatArray; - -/** Device correctness gate for the grouped Q4_0 projection dispatches. */ -public final class Q4GroupedProjectionExperiment { - private static final double MAX_RELATIVE_L2 = 2.0e-5; - private static final int BATCH_SIZE = 3; - private static final int COLS = 96; - private static final int FIRST_ROWS = 11; - private static final int SECOND_ROWS = 7; - private static final int THIRD_ROWS = 5; - - private Q4GroupedProjectionExperiment() {} - - public static void main(String[] args) throws TornadoExecutionPlanException { - byte[] firstWeights = Q4ProjectionExperiment.randomQ4Matrix(FIRST_ROWS, COLS, 41L); - byte[] secondWeights = Q4ProjectionExperiment.randomQ4Matrix(SECOND_ROWS, COLS, 43L); - byte[] thirdWeights = Q4ProjectionExperiment.randomQ4Matrix(THIRD_ROWS, COLS, 47L); - float[] input = Q4ProjectionExperiment.randomFloats(BATCH_SIZE * COLS, 53L); - byte[] activations = new byte[BATCH_SIZE * COLS]; - float[] scales = new float[BATCH_SIZE * COLS / 32]; - Q4ProjectionKernel.quantize(input, activations, scales, BATCH_SIZE, COLS); - - float[] expectedFirst = - Q4ProjectionExperiment.vectorApiProjection( - firstWeights, input, BATCH_SIZE, FIRST_ROWS, COLS); - float[] expectedSecond = - Q4ProjectionExperiment.vectorApiProjection( - secondWeights, input, BATCH_SIZE, SECOND_ROWS, COLS); - float[] expectedThird = - Q4ProjectionExperiment.vectorApiProjection( - thirdWeights, input, BATCH_SIZE, THIRD_ROWS, COLS); - - double[] dualErrors = - runDual(firstWeights, secondWeights, activations, scales, expectedFirst, expectedSecond); - double[] tripleErrors = - runTriple( - firstWeights, - secondWeights, - thirdWeights, - activations, - scales, - expectedFirst, - expectedSecond, - expectedThird); - System.out.printf( - Locale.ROOT, "dual relative L2 first=%.8g second=%.8g%n", dualErrors[0], dualErrors[1]); - System.out.printf( - Locale.ROOT, - "triple relative L2 first=%.8g second=%.8g third=%.8g%n", - tripleErrors[0], - tripleErrors[1], - tripleErrors[2]); - } - - private static double[] runDual( - byte[] firstWeights, - byte[] secondWeights, - byte[] activations, - float[] scales, - float[] expectedFirst, - float[] expectedSecond) - throws TornadoExecutionPlanException { - ByteArray deviceFirstWeights = ByteArray.fromArray(firstWeights); - ByteArray deviceSecondWeights = ByteArray.fromArray(secondWeights); - ByteArray deviceActivations = ByteArray.fromArray(activations); - FloatArray deviceScales = FloatArray.fromArray(scales); - FloatArray firstOutput = new FloatArray(BATCH_SIZE * FIRST_ROWS); - FloatArray secondOutput = new FloatArray(BATCH_SIZE * SECOND_ROWS); - TaskGraph graph = - new TaskGraph("q4-dual-projection") - .transferToDevice( - DataTransferMode.FIRST_EXECUTION, - deviceFirstWeights, - deviceSecondWeights, - deviceActivations, - deviceScales) - .task( - "multiply-dual", - Q4ProjectionKernel::multiplyDual, - deviceFirstWeights, - FIRST_ROWS, - deviceSecondWeights, - SECOND_ROWS, - deviceActivations, - deviceScales, - firstOutput, - secondOutput, - BATCH_SIZE, - COLS) - .transferToHost(DataTransferMode.EVERY_EXECUTION, firstOutput, secondOutput); - try (TornadoExecutionPlan plan = new TornadoExecutionPlan(graph.snapshot())) { - plan.execute(); - } - return checkedErrors( - new float[][] {expectedFirst, expectedSecond}, - new FloatArray[] {firstOutput, secondOutput}, - "dual"); - } - - private static double[] runTriple( - byte[] firstWeights, - byte[] secondWeights, - byte[] thirdWeights, - byte[] activations, - float[] scales, - float[] expectedFirst, - float[] expectedSecond, - float[] expectedThird) - throws TornadoExecutionPlanException { - ByteArray deviceFirstWeights = ByteArray.fromArray(firstWeights); - ByteArray deviceSecondWeights = ByteArray.fromArray(secondWeights); - ByteArray deviceThirdWeights = ByteArray.fromArray(thirdWeights); - ByteArray deviceActivations = ByteArray.fromArray(activations); - FloatArray deviceScales = FloatArray.fromArray(scales); - FloatArray firstOutput = new FloatArray(BATCH_SIZE * FIRST_ROWS); - FloatArray secondOutput = new FloatArray(BATCH_SIZE * SECOND_ROWS); - FloatArray thirdOutput = new FloatArray(BATCH_SIZE * THIRD_ROWS); - TaskGraph graph = - new TaskGraph("q4-triple-projection") - .transferToDevice( - DataTransferMode.FIRST_EXECUTION, - deviceFirstWeights, - deviceSecondWeights, - deviceThirdWeights, - deviceActivations, - deviceScales) - .task( - "multiply-triple", - Q4ProjectionKernel::multiplyTriple, - deviceFirstWeights, - FIRST_ROWS, - deviceSecondWeights, - SECOND_ROWS, - deviceThirdWeights, - THIRD_ROWS, - deviceActivations, - deviceScales, - firstOutput, - secondOutput, - thirdOutput, - BATCH_SIZE, - COLS) - .transferToHost( - DataTransferMode.EVERY_EXECUTION, firstOutput, secondOutput, thirdOutput); - try (TornadoExecutionPlan plan = new TornadoExecutionPlan(graph.snapshot())) { - plan.execute(); - } - return checkedErrors( - new float[][] {expectedFirst, expectedSecond, expectedThird}, - new FloatArray[] {firstOutput, secondOutput, thirdOutput}, - "triple"); - } - - private static double[] checkedErrors( - float[][] expected, FloatArray[] actual, String projectionKind) { - double[] errors = new double[expected.length]; - for (int projection = 0; projection < expected.length; projection++) { - errors[projection] = - Q4ProjectionExperiment.relativeL2(expected[projection], actual[projection].toHeapArray()); - if (errors[projection] > MAX_RELATIVE_L2) { - throw new IllegalStateException( - projectionKind - + " projection " - + projection - + " failed relative-L2 gate: " - + errors[projection]); - } - } - return errors; - } -} diff --git a/models-accelerator-bench/src/main/java/com/integrallis/models/accelerator/Q4ProjectionExperiment.java b/models-accelerator-bench/src/main/java/com/integrallis/models/accelerator/Q4ProjectionExperiment.java deleted file mode 100644 index 0b894694..00000000 --- a/models-accelerator-bench/src/main/java/com/integrallis/models/accelerator/Q4ProjectionExperiment.java +++ /dev/null @@ -1,330 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.accelerator; - -import com.integrallis.models.backend.tornado.Q4ProjectionKernel; -import com.integrallis.vectors.core.GgufQ4Kernel; -import com.integrallis.vectors.core.VectorUtil; -import java.lang.foreign.MemorySegment; -import java.util.Arrays; -import java.util.Locale; -import java.util.SplittableRandom; -import uk.ac.manchester.tornado.api.TaskGraph; -import uk.ac.manchester.tornado.api.TornadoExecutionPlan; -import uk.ac.manchester.tornado.api.TornadoExecutionResult; -import uk.ac.manchester.tornado.api.enums.DataTransferMode; -import uk.ac.manchester.tornado.api.enums.ProfilerMode; -import uk.ac.manchester.tornado.api.exceptions.TornadoExecutionPlanException; -import uk.ac.manchester.tornado.api.types.arrays.ByteArray; -import uk.ac.manchester.tornado.api.types.arrays.FloatArray; - -/** Measures one isolated, production-compatible Java-authored Q4_0 by Q8_0 projection. */ -public final class Q4ProjectionExperiment { - private static final int BLOCK_VALUES = 32; - private static final int BLOCK_BYTES = 18; - - private Q4ProjectionExperiment() {} - - public static void main(String[] args) throws TornadoExecutionPlanException { - Configuration configuration = Configuration.parse(args); - int activationEntries = Math.multiplyExact(configuration.batchSize(), configuration.cols()); - int scaleEntries = activationEntries / BLOCK_VALUES; - byte[] weights = randomQ4Matrix(configuration.rows(), configuration.cols(), 31L); - float[] input = randomFloats(activationEntries, 37L); - float[] expected = - vectorApiProjection( - weights, input, configuration.batchSize(), configuration.rows(), configuration.cols()); - byte[] preparedActivations = new byte[activationEntries]; - float[] preparedScales = new float[scaleEntries]; - Q4ProjectionKernel.quantize( - input, - preparedActivations, - preparedScales, - configuration.batchSize(), - configuration.cols()); - - ByteArray deviceWeights = ByteArray.fromArray(weights); - ByteArray deviceActivations = ByteArray.fromArray(preparedActivations); - FloatArray deviceScales = FloatArray.fromArray(preparedScales); - FloatArray deviceOutput = - new FloatArray(Math.multiplyExact(configuration.batchSize(), configuration.rows())); - Q4ProjectionKernel.validate( - deviceWeights, - deviceActivations, - deviceScales, - deviceOutput, - configuration.batchSize(), - configuration.rows(), - configuration.cols()); - - TaskGraph graph = - new TaskGraph("q4-projection") - .transferToDevice(DataTransferMode.FIRST_EXECUTION, deviceWeights) - .transferToDevice(DataTransferMode.EVERY_EXECUTION, deviceActivations, deviceScales) - .task( - "multiply", - Q4ProjectionKernel::multiply, - deviceWeights, - deviceActivations, - deviceScales, - deviceOutput, - configuration.batchSize(), - configuration.rows(), - configuration.cols()) - .transferToHost(DataTransferMode.EVERY_EXECUTION, deviceOutput); - - long coldNanos; - long[] warmNanos = new long[configuration.iterations()]; - long[] preparationNanos = new long[configuration.iterations()]; - long[] endToEndNanos = new long[configuration.iterations()]; - TornadoExecutionResult lastResult; - try (TornadoExecutionPlan plan = - new TornadoExecutionPlan(graph.snapshot()).withProfiler(ProfilerMode.SILENT)) { - long started = System.nanoTime(); - lastResult = plan.execute(); - coldNanos = System.nanoTime() - started; - for (int warmup = 0; warmup < configuration.warmups(); warmup++) { - prepareAndStage( - input, - preparedActivations, - preparedScales, - deviceActivations, - deviceScales, - configuration.batchSize(), - configuration.cols()); - plan.execute(); - } - for (int iteration = 0; iteration < warmNanos.length; iteration++) { - long endToEndStarted = System.nanoTime(); - started = System.nanoTime(); - prepareAndStage( - input, - preparedActivations, - preparedScales, - deviceActivations, - deviceScales, - configuration.batchSize(), - configuration.cols()); - preparationNanos[iteration] = System.nanoTime() - started; - started = System.nanoTime(); - lastResult = plan.execute(); - warmNanos[iteration] = System.nanoTime() - started; - endToEndNanos[iteration] = System.nanoTime() - endToEndStarted; - } - } - - float[] actual = deviceOutput.toHeapArray(); - double relativeL2 = relativeL2(expected, actual); - if (relativeL2 > 2.0e-5) { - throw new IllegalStateException( - "accelerator result failed relative-L2 gate: " - + relativeL2 - + ", expected[0]=" - + expected[0] - + ", actual[0]=" - + actual[0]); - } - - long[] cpuNanos = - timeCpu( - weights, - input, - configuration.batchSize(), - configuration.rows(), - configuration.cols(), - configuration.warmups(), - configuration.iterations()); - long deviceMedian = median(warmNanos); - long preparationMedian = median(preparationNanos); - long endToEndMedian = median(endToEndNanos); - long cpuMedian = median(cpuNanos); - - System.out.printf( - Locale.ROOT, - "shape batch=%d rows=%d cols=%d%n", - configuration.batchSize(), - configuration.rows(), - configuration.cols()); - System.out.printf( - Locale.ROOT, - "weights %.2f MiB, persistent after first execution%n", - weights.length / (1024.0 * 1024.0)); - System.out.printf(Locale.ROOT, "correctness relative L2 %.8g%n", relativeL2); - System.out.printf(Locale.ROOT, "cold device %.3f ms%n", coldNanos / 1_000_000.0); - System.out.printf(Locale.ROOT, "host Q8 staging %.3f ms%n", preparationMedian / 1_000_000.0); - System.out.printf(Locale.ROOT, "warm device p50 %.3f ms%n", deviceMedian / 1_000_000.0); - System.out.printf(Locale.ROOT, "end-to-end p50 %.3f ms%n", endToEndMedian / 1_000_000.0); - System.out.printf(Locale.ROOT, "Vector API p50 %.3f ms%n", cpuMedian / 1_000_000.0); - System.out.printf( - Locale.ROOT, "speedup %.2fx%n", (double) cpuMedian / endToEndMedian); - System.out.printf( - Locale.ROOT, - "last kernel %.3f ms%n", - lastResult.getProfilerResult().getDeviceKernelTime() / 1_000_000.0); - System.out.printf( - Locale.ROOT, - "last transfers %.3f ms%n", - lastResult.getProfilerResult().getDataTransfersTime() / 1_000_000.0); - } - - private static void prepareAndStage( - float[] input, - byte[] preparedActivations, - float[] preparedScales, - ByteArray deviceActivations, - FloatArray deviceScales, - int batchSize, - int cols) { - Q4ProjectionKernel.quantize(input, preparedActivations, preparedScales, batchSize, cols); - deviceActivations.getSegment().copyFrom(MemorySegment.ofArray(preparedActivations)); - deviceScales.getSegment().copyFrom(MemorySegment.ofArray(preparedScales)); - } - - private static long[] timeCpu( - byte[] weights, - float[] input, - int batchSize, - int rows, - int cols, - int warmups, - int iterations) { - MemorySegment weightSegment = MemorySegment.ofArray(weights); - float[] output = new float[Math.multiplyExact(batchSize, rows)]; - byte[] quants = new byte[Math.multiplyExact(batchSize, cols)]; - float[] scales = new float[Math.multiplyExact(batchSize, cols / BLOCK_VALUES)]; - int[] corrections = new int[Math.multiplyExact(batchSize, (cols + 3) / 4)]; - float[] lanes = new float[Math.multiplyExact(output.length, 8)]; - for (int warmup = 0; warmup < warmups; warmup++) { - VectorUtil.ggufQ4_0Q8_0BatchedMatmul( - input, - weightSegment, - batchSize, - rows, - cols, - output, - quants, - scales, - corrections, - lanes, - GgufQ4Kernel.WIDENED); - } - long[] nanos = new long[iterations]; - for (int iteration = 0; iteration < iterations; iteration++) { - long started = System.nanoTime(); - VectorUtil.ggufQ4_0Q8_0BatchedMatmul( - input, - weightSegment, - batchSize, - rows, - cols, - output, - quants, - scales, - corrections, - lanes, - GgufQ4Kernel.WIDENED); - nanos[iteration] = System.nanoTime() - started; - } - return nanos; - } - - static float[] vectorApiProjection( - byte[] weights, float[] input, int batchSize, int rows, int cols) { - float[] output = new float[Math.multiplyExact(batchSize, rows)]; - VectorUtil.ggufQ4_0Q8_0BatchedMatmul( - input, - MemorySegment.ofArray(weights), - batchSize, - rows, - cols, - output, - new byte[Math.multiplyExact(batchSize, cols)], - new float[Math.multiplyExact(batchSize, cols / BLOCK_VALUES)], - new int[Math.multiplyExact(batchSize, (cols + 3) / 4)], - new float[Math.multiplyExact(output.length, 8)], - GgufQ4Kernel.WIDENED); - return output; - } - - static double relativeL2(float[] expected, float[] actual) { - double squaredError = 0.0; - double squaredReference = 0.0; - for (int index = 0; index < expected.length; index++) { - double difference = actual[index] - expected[index]; - squaredError += difference * difference; - squaredReference += expected[index] * expected[index]; - } - return Math.sqrt(squaredError / squaredReference); - } - - private static long median(long[] samples) { - long[] sorted = samples.clone(); - Arrays.sort(sorted); - return sorted[sorted.length / 2]; - } - - static byte[] randomQ4Matrix(int rows, int cols, long seed) { - SplittableRandom random = new SplittableRandom(seed); - int blocks = Math.multiplyExact(rows, cols / BLOCK_VALUES); - byte[] weights = new byte[Math.multiplyExact(blocks, BLOCK_BYTES)]; - for (int block = 0; block < blocks; block++) { - short scale = Float.floatToFloat16(0.001f + random.nextFloat() * 0.05f); - int offset = block * BLOCK_BYTES; - weights[offset] = (byte) scale; - weights[offset + 1] = (byte) (scale >>> 8); - for (int quant = 0; quant < 16; quant++) { - weights[offset + 2 + quant] = (byte) random.nextInt(256); - } - } - return weights; - } - - static float[] randomFloats(int length, long seed) { - SplittableRandom random = new SplittableRandom(seed); - float[] values = new float[length]; - for (int index = 0; index < length; index++) { - values[index] = random.nextFloat(-2.0f, 2.0f); - } - return values; - } - - private record Configuration(int batchSize, int rows, int cols, int warmups, int iterations) { - private static Configuration parse(String[] args) { - if (args.length != 0 && args.length != 5) { - throw new IllegalArgumentException( - "usage: Q4ProjectionExperiment [batch rows cols warmups iterations]"); - } - Configuration configuration = - args.length == 0 - ? new Configuration(32, 3072, 1024, 3, 10) - : new Configuration( - Integer.parseInt(args[0]), - Integer.parseInt(args[1]), - Integer.parseInt(args[2]), - Integer.parseInt(args[3]), - Integer.parseInt(args[4])); - if (configuration.batchSize < 1 - || configuration.rows < 1 - || configuration.cols < BLOCK_VALUES - || configuration.cols % BLOCK_VALUES != 0 - || configuration.warmups < 0 - || configuration.iterations < 1) { - throw new IllegalArgumentException("invalid projection experiment configuration"); - } - return configuration; - } - } -} diff --git a/models-accelerator-bench/src/main/java/com/integrallis/models/accelerator/Q8ProjectionExperiment.java b/models-accelerator-bench/src/main/java/com/integrallis/models/accelerator/Q8ProjectionExperiment.java deleted file mode 100644 index f6441dc4..00000000 --- a/models-accelerator-bench/src/main/java/com/integrallis/models/accelerator/Q8ProjectionExperiment.java +++ /dev/null @@ -1,241 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.accelerator; - -import com.integrallis.vectors.core.VectorUtil; -import java.lang.foreign.MemorySegment; -import java.util.Arrays; -import java.util.Locale; -import java.util.SplittableRandom; -import uk.ac.manchester.tornado.api.TaskGraph; -import uk.ac.manchester.tornado.api.TornadoExecutionPlan; -import uk.ac.manchester.tornado.api.TornadoExecutionResult; -import uk.ac.manchester.tornado.api.enums.DataTransferMode; -import uk.ac.manchester.tornado.api.enums.ProfilerMode; -import uk.ac.manchester.tornado.api.exceptions.TornadoExecutionPlanException; -import uk.ac.manchester.tornado.api.types.arrays.ByteArray; -import uk.ac.manchester.tornado.api.types.arrays.FloatArray; - -/** Measures one isolated Java-authored Q8_0 projection on a TornadoVM device. */ -public final class Q8ProjectionExperiment { - private static final int BLOCK_VALUES = 32; - private static final int BLOCK_BYTES = 34; - - private Q8ProjectionExperiment() {} - - public static void main(String[] args) throws TornadoExecutionPlanException { - Configuration configuration = Configuration.parse(args); - byte[] weights = randomQ8Matrix(configuration.rows(), configuration.cols(), 23L); - float[] input = - randomFloats(Math.multiplyExact(configuration.batchSize(), configuration.cols()), 29L); - float[] expected = - vectorApiProjection( - weights, input, configuration.batchSize(), configuration.rows(), configuration.cols()); - - ByteArray deviceWeights = ByteArray.fromArray(weights); - FloatArray deviceInput = FloatArray.fromArray(input); - FloatArray deviceOutput = - new FloatArray(Math.multiplyExact(configuration.batchSize(), configuration.rows())); - Q8ProjectionKernel.validate( - deviceWeights, - deviceInput, - deviceOutput, - configuration.batchSize(), - configuration.rows(), - configuration.cols()); - - TaskGraph graph = - new TaskGraph("q8-projection") - .transferToDevice(DataTransferMode.FIRST_EXECUTION, deviceWeights) - .transferToDevice(DataTransferMode.EVERY_EXECUTION, deviceInput) - .task( - "multiply", - Q8ProjectionKernel::multiply, - deviceWeights, - deviceInput, - deviceOutput, - configuration.batchSize(), - configuration.rows(), - configuration.cols()) - .transferToHost(DataTransferMode.EVERY_EXECUTION, deviceOutput); - - long coldNanos; - long[] warmNanos = new long[configuration.iterations()]; - TornadoExecutionResult lastResult; - try (TornadoExecutionPlan plan = - new TornadoExecutionPlan(graph.snapshot()).withProfiler(ProfilerMode.SILENT)) { - long started = System.nanoTime(); - lastResult = plan.execute(); - coldNanos = System.nanoTime() - started; - for (int warmup = 0; warmup < configuration.warmups(); warmup++) { - plan.execute(); - } - for (int iteration = 0; iteration < warmNanos.length; iteration++) { - started = System.nanoTime(); - lastResult = plan.execute(); - warmNanos[iteration] = System.nanoTime() - started; - } - } - - float[] actual = deviceOutput.toHeapArray(); - double relativeL2 = relativeL2(expected, actual); - if (relativeL2 > 2.0e-5) { - throw new IllegalStateException("accelerator result failed relative-L2 gate: " + relativeL2); - } - - long[] cpuNanos = - timeCpu( - weights, - input, - configuration.batchSize(), - configuration.rows(), - configuration.cols(), - configuration.warmups(), - configuration.iterations()); - long deviceMedian = median(warmNanos); - long cpuMedian = median(cpuNanos); - double weightGiB = weights.length / (1024.0 * 1024.0 * 1024.0); - double effectiveGiBPerSecond = - weightGiB * configuration.batchSize() / (deviceMedian / 1_000_000_000.0); - - System.out.printf( - Locale.ROOT, - "shape batch=%d rows=%d cols=%d%n", - configuration.batchSize(), - configuration.rows(), - configuration.cols()); - System.out.printf( - Locale.ROOT, - "weights %.2f MiB, persistent after first execution%n", - weights.length / (1024.0 * 1024.0)); - System.out.printf(Locale.ROOT, "correctness relative L2 %.8g%n", relativeL2); - System.out.printf(Locale.ROOT, "cold device %.3f ms%n", coldNanos / 1_000_000.0); - System.out.printf(Locale.ROOT, "warm device p50 %.3f ms%n", deviceMedian / 1_000_000.0); - System.out.printf(Locale.ROOT, "Vector API p50 %.3f ms%n", cpuMedian / 1_000_000.0); - System.out.printf(Locale.ROOT, "speedup %.2fx%n", (double) cpuMedian / deviceMedian); - System.out.printf(Locale.ROOT, "effective traffic %.2f GiB/s%n", effectiveGiBPerSecond); - System.out.printf( - Locale.ROOT, - "last kernel %.3f ms%n", - lastResult.getProfilerResult().getDeviceKernelTime() / 1_000_000.0); - System.out.printf( - Locale.ROOT, - "last transfers %.3f ms%n", - lastResult.getProfilerResult().getDataTransfersTime() / 1_000_000.0); - } - - private static long[] timeCpu( - byte[] weights, - float[] input, - int batchSize, - int rows, - int cols, - int warmups, - int iterations) { - for (int warmup = 0; warmup < warmups; warmup++) { - vectorApiProjection(weights, input, batchSize, rows, cols); - } - long[] nanos = new long[iterations]; - for (int iteration = 0; iteration < iterations; iteration++) { - long started = System.nanoTime(); - vectorApiProjection(weights, input, batchSize, rows, cols); - nanos[iteration] = System.nanoTime() - started; - } - return nanos; - } - - private static float[] vectorApiProjection( - byte[] weights, float[] input, int batchSize, int rows, int cols) { - float[] output = new float[Math.multiplyExact(batchSize, rows)]; - MemorySegment weightSegment = MemorySegment.ofArray(weights); - for (int batch = 0; batch < batchSize; batch++) { - float[] query = Arrays.copyOfRange(input, batch * cols, (batch + 1) * cols); - float[] projected = new float[rows]; - VectorUtil.ggufQ8_0BatchDotProduct(query, weightSegment, rows, cols, projected); - System.arraycopy(projected, 0, output, batch * rows, rows); - } - return output; - } - - private static double relativeL2(float[] expected, float[] actual) { - double squaredError = 0.0; - double squaredReference = 0.0; - for (int index = 0; index < expected.length; index++) { - double difference = actual[index] - expected[index]; - squaredError += difference * difference; - squaredReference += expected[index] * expected[index]; - } - return Math.sqrt(squaredError / squaredReference); - } - - private static long median(long[] samples) { - long[] sorted = samples.clone(); - Arrays.sort(sorted); - return sorted[sorted.length / 2]; - } - - private static byte[] randomQ8Matrix(int rows, int cols, long seed) { - SplittableRandom random = new SplittableRandom(seed); - int blocks = Math.multiplyExact(rows, cols / BLOCK_VALUES); - byte[] weights = new byte[Math.multiplyExact(blocks, BLOCK_BYTES)]; - for (int block = 0; block < blocks; block++) { - short scale = Float.floatToFloat16(0.001f + random.nextFloat() * 0.05f); - int offset = block * BLOCK_BYTES; - weights[offset] = (byte) scale; - weights[offset + 1] = (byte) (scale >>> 8); - for (int quant = 0; quant < BLOCK_VALUES; quant++) { - weights[offset + 2 + quant] = (byte) (random.nextInt(255) - 127); - } - } - return weights; - } - - private static float[] randomFloats(int length, long seed) { - SplittableRandom random = new SplittableRandom(seed); - float[] values = new float[length]; - for (int index = 0; index < length; index++) { - values[index] = random.nextFloat(-2.0f, 2.0f); - } - return values; - } - - private record Configuration(int batchSize, int rows, int cols, int warmups, int iterations) { - private static Configuration parse(String[] args) { - if (args.length != 0 && args.length != 5) { - throw new IllegalArgumentException( - "usage: Q8ProjectionExperiment [batch rows cols warmups iterations]"); - } - Configuration configuration = - args.length == 0 - ? new Configuration(32, 3072, 1024, 3, 10) - : new Configuration( - Integer.parseInt(args[0]), - Integer.parseInt(args[1]), - Integer.parseInt(args[2]), - Integer.parseInt(args[3]), - Integer.parseInt(args[4])); - if (configuration.batchSize < 1 - || configuration.rows < 1 - || configuration.cols < BLOCK_VALUES - || configuration.cols % BLOCK_VALUES != 0 - || configuration.warmups < 0 - || configuration.iterations < 1) { - throw new IllegalArgumentException("invalid projection experiment configuration"); - } - return configuration; - } - } -} diff --git a/models-accelerator-bench/src/main/java/com/integrallis/models/accelerator/Q8ProjectionKernel.java b/models-accelerator-bench/src/main/java/com/integrallis/models/accelerator/Q8ProjectionKernel.java deleted file mode 100644 index 0c4dd4a4..00000000 --- a/models-accelerator-bench/src/main/java/com/integrallis/models/accelerator/Q8ProjectionKernel.java +++ /dev/null @@ -1,73 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.accelerator; - -import uk.ac.manchester.tornado.api.annotations.Parallel; -import uk.ac.manchester.tornado.api.types.arrays.ByteArray; -import uk.ac.manchester.tornado.api.types.arrays.FloatArray; - -/** Java-authored TornadoVM experiment for row-major GGUF Q8_0 projections. */ -final class Q8ProjectionKernel { - private static final int BLOCK_VALUES = 32; - private static final int BLOCK_BYTES = 34; - - private Q8ProjectionKernel() {} - - /** Runs one work item per batch/output-row pair. */ - static void multiply( - ByteArray weights, FloatArray input, FloatArray output, int batchSize, int rows, int cols) { - for (@Parallel int outputIndex = 0; outputIndex < batchSize * rows; outputIndex++) { - int batch = outputIndex / rows; - int row = outputIndex - batch * rows; - int rowBlockOffset = row * (cols / BLOCK_VALUES); - int inputOffset = batch * cols; - float sum = 0.0f; - for (int col = 0; col < cols; col++) { - int block = col / BLOCK_VALUES; - int withinBlock = col - block * BLOCK_VALUES; - int blockByteOffset = (rowBlockOffset + block) * BLOCK_BYTES; - float scale = weights.getHalfFloat(blockByteOffset).getFloat32(); - byte quantized = weights.get(blockByteOffset + 2 + withinBlock); - sum += ((float) quantized * scale) * input.get(inputOffset + col); - } - output.set(outputIndex, sum); - } - } - - static void validate( - ByteArray weights, FloatArray input, FloatArray output, int batchSize, int rows, int cols) { - if (batchSize < 1) { - throw new IllegalArgumentException("batchSize must be positive"); - } - if (rows < 1) { - throw new IllegalArgumentException("rows must be positive"); - } - if (cols < BLOCK_VALUES || cols % BLOCK_VALUES != 0) { - throw new IllegalArgumentException("cols must be a positive multiple of 32"); - } - int expectedWeightBytes = - Math.multiplyExact(Math.multiplyExact(rows, cols / BLOCK_VALUES), BLOCK_BYTES); - if (weights.getSize() != expectedWeightBytes) { - throw new IllegalArgumentException("weights do not match the projection shape"); - } - if (input.getSize() != Math.multiplyExact(batchSize, cols)) { - throw new IllegalArgumentException("input does not match the projection shape"); - } - if (output.getSize() != Math.multiplyExact(batchSize, rows)) { - throw new IllegalArgumentException("output does not match the projection shape"); - } - } -} diff --git a/models-accelerator-bench/src/main/java/com/integrallis/models/accelerator/QwenFullModelExperiment.java b/models-accelerator-bench/src/main/java/com/integrallis/models/accelerator/QwenFullModelExperiment.java deleted file mode 100644 index 2a6f16ac..00000000 --- a/models-accelerator-bench/src/main/java/com/integrallis/models/accelerator/QwenFullModelExperiment.java +++ /dev/null @@ -1,236 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.accelerator; - -import com.integrallis.models.api.ModelPrompt; -import com.integrallis.models.api.SamplingOptions; -import com.integrallis.models.backend.purejava.PureJavaBackend; -import com.integrallis.models.backend.purejava.spi.BatchedCausalAttentionKernel; -import com.integrallis.models.backend.tornado.TornadoGgufBatchedMatrixKernel; -import com.integrallis.models.runtime.GenerationLoop; -import com.integrallis.models.runtime.GenerationMetrics; -import com.integrallis.models.runtime.chat.ChatMessage; -import com.integrallis.models.runtime.chat.ChatTemplate; -import java.nio.file.Files; -import java.nio.file.Path; -import java.time.Duration; -import java.util.Arrays; -import java.util.List; -import java.util.Locale; - -/** Compares a complete Qwen generation on the Vector API and experimental Java/Tornado prefill. */ -public final class QwenFullModelExperiment { - private static final ModelPrompt SHORT_PROMPT = - ChatTemplate.CHATML_NO_THINK.render( - List.of( - ChatMessage.system("Follow the user's output format exactly."), - ChatMessage.user("Reply with exactly: JAVA"))); - private static final String TRANSIT_CONTEXT = - "Route Blue serves Central Station, Museum Square, River Market, and Airport Terminal. " - + "Trains leave Central Station every twelve minutes from 06:00 through 23:00. " - + "River Market and Airport Terminal have step-free platforms and working elevators. "; - private static final SamplingOptions SAMPLING = - SamplingOptions.builder().temperature(0.0f).maxTokens(4).build(); - - private QwenFullModelExperiment() {} - - public static void main(String[] args) { - if (args.length < 1 - || args.length > 6 - || Arrays.stream(args, 1, args.length) - .anyMatch( - argument -> - !"--cpu-only".equals(argument) - && !"--long-prompt".equals(argument) - && !"--eager".equals(argument) - && !"--attention".equals(argument) - && !"--decode".equals(argument))) { - throw new IllegalArgumentException( - "usage: QwenFullModelExperiment [--cpu-only] [--long-prompt] [--eager] [--attention] [--decode]"); - } - boolean cpuOnly = Arrays.asList(args).contains("--cpu-only"); - boolean longPrompt = Arrays.asList(args).contains("--long-prompt"); - boolean eager = Arrays.asList(args).contains("--eager"); - boolean attention = Arrays.asList(args).contains("--attention"); - boolean decode = Arrays.asList(args).contains("--decode"); - if (cpuOnly && (eager || attention || decode)) { - throw new IllegalArgumentException( - "--eager, --attention, and --decode apply only to the Tornado GPU experiment"); - } - ModelPrompt prompt = longPrompt ? longPrompt() : SHORT_PROMPT; - Path model = Path.of(args[0]).toAbsolutePath().normalize(); - if (!Files.isRegularFile(model)) { - throw new IllegalArgumentException("model is not a regular file: " + model); - } - System.setProperty(PureJavaBackend.MAX_CONTEXT_LENGTH_PROPERTY, "512"); - - DecodeProbe cpuDecode = null; - try (PureJavaBackend backend = PureJavaBackend.load(model)) { - run("Vector API cold", backend, prompt); - run("Vector API warm", backend, prompt); - if (decode) { - cpuDecode = runDecodeProbe("Vector API", backend, prompt, 8); - } - } - if (cpuOnly) { - return; - } - - TornadoGgufBatchedMatrixKernel kernel = new TornadoGgufBatchedMatrixKernel(32, decode); - BatchedCausalAttentionKernel attentionKernel = - attention ? new TornadoBatchedCausalAttentionKernel() : BatchedCausalAttentionKernel.none(); - try (PureJavaBackend backend = PureJavaBackend.load(model, kernel, attentionKernel)) { - if (eager) { - prepare( - backend, - kernel, - attention ? (TornadoBatchedCausalAttentionKernel) attentionKernel : null, - prompt); - } - run(eager ? "Tornado first visible" : "Tornado cold", backend, prompt); - run("Tornado warm", backend, prompt); - if (decode) { - DecodeProbe gpuDecode = runDecodeProbe("Tornado", backend, prompt, 8); - if (!Arrays.equals(cpuDecode.tokens(), gpuDecode.tokens())) { - throw new IllegalStateException( - "accelerated decode changed the greedy token sequence: " - + Arrays.toString(cpuDecode.tokens()) - + " != " - + Arrays.toString(gpuDecode.tokens())); - } - } - System.out.printf( - Locale.ROOT, - "accelerated calls=%d plans=%d cumulative=%.3f s%n", - kernel.calls(), - kernel.projectionPlanCount(), - kernel.totalMillis() / 1_000.0); - if (attentionKernel instanceof TornadoBatchedCausalAttentionKernel tornadoAttention) { - System.out.printf( - Locale.ROOT, - "attention calls=%d plans=%d cumulative=%.3f s%n", - tornadoAttention.calls(), - tornadoAttention.planCount(), - tornadoAttention.totalMillis() / 1_000.0); - } - } - } - - private static ModelPrompt longPrompt() { - return ChatTemplate.CHATML_NO_THINK.render( - List.of( - ChatMessage.system("Use the supplied context and follow the output format exactly."), - ChatMessage.user( - "Context: " - + TRANSIT_CONTEXT.repeat(4) - + "Validation instruction: reply with exactly: JAVA"))); - } - - private static void run(String label, PureJavaBackend backend, ModelPrompt prompt) { - GenerationLoop loop = new GenerationLoop(backend); - String output = loop.generate(prompt, SAMPLING); - GenerationMetrics metrics = loop.lastGenerationMetrics(); - System.out.printf( - Locale.ROOT, - "%s output=%s promptTokens=%d completionTokens=%d prefill=%.3f s ttft=%.3f s total=%.3f s%n", - label, - output.replace('\n', ' '), - metrics.usage().promptTokens(), - metrics.usage().completionTokens(), - seconds(metrics.prefill()), - seconds(metrics.timeToFirstToken().orElse(Duration.ZERO)), - seconds(metrics.total())); - } - - private static DecodeProbe runDecodeProbe( - String label, PureJavaBackend backend, ModelPrompt prompt, int steps) { - backend.reset(); - int[] promptTokens = backend.tokenizer().encode(prompt); - float[] logits = backend.prefill(promptTokens, 0); - int[] generated = new int[steps]; - long[] nanos = new long[steps]; - for (int step = 0; step < steps; step++) { - int token = argmax(logits); - generated[step] = token; - long started = System.nanoTime(); - logits = backend.forward(token, promptTokens.length + step); - nanos[step] = System.nanoTime() - started; - } - long total = Arrays.stream(nanos).sum(); - long[] sorted = nanos.clone(); - Arrays.sort(sorted); - System.out.printf( - Locale.ROOT, - "%s decode output=%s tokens=%d p50=%.3f ms total=%.3f ms%n", - label, - backend.tokenizer().decode(generated).replace('\n', ' '), - steps, - sorted[sorted.length / 2] / 1_000_000.0, - total / 1_000_000.0); - backend.reset(); - return new DecodeProbe(generated); - } - - private static int argmax(float[] values) { - int maximumIndex = 0; - for (int index = 1; index < values.length; index++) { - if (values[index] > values[maximumIndex]) { - maximumIndex = index; - } - } - return maximumIndex; - } - - private static void prepare( - PureJavaBackend backend, - TornadoGgufBatchedMatrixKernel kernel, - TornadoBatchedCausalAttentionKernel attentionKernel, - ModelPrompt prompt) { - long started = System.nanoTime(); - int[] promptTokens = backend.tokenizer().encode(prompt); - int[] readinessTokens = readinessTokens(promptTokens, kernel.executionBatchSize()); - backend.prefill(readinessTokens, 0); - if (kernel.acceleratesDecode()) { - backend.forward(readinessTokens[0], readinessTokens.length); - } - backend.reset(); - System.out.printf( - Locale.ROOT, - "Tornado readiness=%.3f s projectionPlans=%d projectionCalls=%d attentionPlans=%d attentionCalls=%d%n", - (System.nanoTime() - started) / 1_000_000_000.0, - kernel.projectionPlanCount(), - kernel.calls(), - attentionKernel == null ? 0 : attentionKernel.planCount(), - attentionKernel == null ? 0 : attentionKernel.calls()); - } - - static int[] readinessTokens(int[] source, int executionBatchSize) { - if (source.length == 0) { - throw new IllegalArgumentException("readiness requires prompt tokens"); - } - int[] tokens = new int[executionBatchSize]; - for (int index = 0; index < tokens.length; index++) { - tokens[index] = source[index % source.length]; - } - return tokens; - } - - private static double seconds(Duration duration) { - return duration.toNanos() / 1_000_000_000.0; - } - - private record DecodeProbe(int[] tokens) {} -} diff --git a/models-accelerator-bench/src/main/java/com/integrallis/models/accelerator/TornadoBackendExperiment.java b/models-accelerator-bench/src/main/java/com/integrallis/models/accelerator/TornadoBackendExperiment.java deleted file mode 100644 index 074f8a81..00000000 --- a/models-accelerator-bench/src/main/java/com/integrallis/models/accelerator/TornadoBackendExperiment.java +++ /dev/null @@ -1,107 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.accelerator; - -import com.integrallis.models.api.ModelPrompt; -import com.integrallis.models.api.SamplingOptions; -import com.integrallis.models.backend.purejava.PureJavaBackend; -import com.integrallis.models.backend.tornado.TornadoBackend; -import com.integrallis.models.backend.tornado.TornadoBackendOptions; -import com.integrallis.models.backend.tornado.TornadoBackendRuntime; -import com.integrallis.models.runtime.GenerationLoop; -import com.integrallis.models.runtime.GenerationMetrics; -import com.integrallis.models.runtime.chat.ChatMessage; -import com.integrallis.models.runtime.chat.ChatTemplate; -import java.nio.file.Files; -import java.nio.file.Path; -import java.util.List; -import java.util.Locale; - -/** End-to-end release gate for the public automatic accelerator loader and CPU parity. */ -public final class TornadoBackendExperiment { - private static final ModelPrompt PROMPT = - ChatTemplate.CHATML_NO_THINK.render( - List.of( - ChatMessage.system("Follow the user's output format exactly."), - ChatMessage.user("Reply with exactly: JAVA"))); - private static final SamplingOptions SAMPLING = - SamplingOptions.builder().temperature(0.0f).maxTokens(4).build(); - - private TornadoBackendExperiment() {} - - public static void main(String[] args) { - if (args.length != 1) { - throw new IllegalArgumentException("usage: TornadoBackendExperiment "); - } - Path model = Path.of(args[0]).toAbsolutePath().normalize(); - if (!Files.isRegularFile(model)) { - throw new IllegalArgumentException("model is not a regular file: " + model); - } - System.setProperty(PureJavaBackend.MAX_CONTEXT_LENGTH_PROPERTY, "512"); - String expected; - try (PureJavaBackend backend = PureJavaBackend.load(model)) { - expected = generate("Vector API", backend); - } - TornadoBackendOptions required = new TornadoBackendOptions(true, true, true, 32); - try (TornadoBackendRuntime runtime = TornadoBackend.open(model, required)) { - String actual = generate("Automatic accelerator", runtime.backend()); - if (!expected.equals(actual)) { - throw new IllegalStateException( - "automatic accelerator changed generated text: " + expected + " != " + actual); - } - System.out.printf( - Locale.ROOT, - "selected=%s device=%s readiness=%.3f s reason=%s%n", - runtime.status().accelerated(), - runtime.status().device(), - runtime.status().readinessTime().toNanos() / 1_000_000_000.0, - runtime.status().reason()); - } - try (PureJavaBackend backend = PureJavaBackend.loadAutomatic(model)) { - String actual = generate("Service-loaded accelerator", backend); - if (!expected.equals(actual)) { - throw new IllegalStateException( - "service-loaded accelerator changed generated text: " + expected + " != " + actual); - } - String implementation = - backend - .diagnostics() - .optimization("grouped-projections") - .orElseThrow() - .settings() - .get("implementation"); - if (!"tornadovm-java-q4-prefill-decode".equals(implementation)) { - throw new IllegalStateException("automatic provider was not selected: " + implementation); - } - } - } - - private static String generate(String label, PureJavaBackend backend) { - GenerationLoop loop = new GenerationLoop(backend); - String output = loop.generate(PROMPT, SAMPLING); - GenerationMetrics metrics = loop.lastGenerationMetrics(); - System.out.printf( - Locale.ROOT, - "%s output=%s promptTokens=%d completionTokens=%d prefill=%.3f s total=%.3f s%n", - label, - output.replace('\n', ' '), - metrics.usage().promptTokens(), - metrics.usage().completionTokens(), - metrics.prefill().toNanos() / 1_000_000_000.0, - metrics.total().toNanos() / 1_000_000_000.0); - return output; - } -} diff --git a/models-accelerator-bench/src/main/java/com/integrallis/models/accelerator/TornadoBatchedCausalAttentionKernel.java b/models-accelerator-bench/src/main/java/com/integrallis/models/accelerator/TornadoBatchedCausalAttentionKernel.java deleted file mode 100644 index 2a33049a..00000000 --- a/models-accelerator-bench/src/main/java/com/integrallis/models/accelerator/TornadoBatchedCausalAttentionKernel.java +++ /dev/null @@ -1,178 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.accelerator; - -import com.integrallis.models.backend.purejava.spi.BatchedCausalAttentionKernel; -import java.util.LinkedHashMap; -import java.util.Map; - -/** Experimental Java/Tornado batched attention provider with one retained KV cache per layer. */ -public final class TornadoBatchedCausalAttentionKernel implements BatchedCausalAttentionKernel { - private static final int DEFAULT_EXECUTION_BATCH_SIZE = 32; - private static final int MINIMUM_BATCH_SIZE = 4; - - private final int executionBatchSize; - private final Map plans = new LinkedHashMap<>(); - private final Map nextPositions = new LinkedHashMap<>(); - private long calls; - private long totalNanos; - private boolean closed; - - public TornadoBatchedCausalAttentionKernel() { - this(DEFAULT_EXECUTION_BATCH_SIZE); - } - - public TornadoBatchedCausalAttentionKernel(int executionBatchSize) { - if (executionBatchSize < MINIMUM_BATCH_SIZE) { - throw new IllegalArgumentException( - "executionBatchSize must be at least " + MINIMUM_BATCH_SIZE); - } - this.executionBatchSize = executionBatchSize; - } - - @Override - public synchronized boolean isEligible( - int layer, - int startPosition, - int batchSize, - int numHeads, - int numKvHeads, - int keyLength, - int valueLength, - int maxSequenceLength, - int slidingWindow) { - return !closed - && layer >= 0 - && startPosition >= 0 - && batchSize >= MINIMUM_BATCH_SIZE - && batchSize <= executionBatchSize - && numHeads > 0 - && numKvHeads > 0 - && numHeads % numKvHeads == 0 - && keyLength > 0 - && valueLength == keyLength - && maxSequenceLength >= startPosition + batchSize - && slidingWindow == 0 - && nextPositions.getOrDefault(layer, 0) == startPosition; - } - - @Override - public synchronized void attend( - float[] output, - float[] query, - float[] key, - float[] value, - int layer, - int startPosition, - int batchSize, - int numHeads, - int numKvHeads, - int keyLength, - int valueLength, - int maxSequenceLength, - int slidingWindow) { - if (!isEligible( - layer, - startPosition, - batchSize, - numHeads, - numKvHeads, - keyLength, - valueLength, - maxSequenceLength, - slidingWindow)) { - throw new UnsupportedOperationException( - "attention chunk is not eligible for the Tornado experiment"); - } - AttentionKey planKey = - new AttentionKey( - layer, numHeads, numKvHeads, keyLength, valueLength, maxSequenceLength, slidingWindow); - TornadoCausalAttentionPlan plan = - plans.computeIfAbsent( - planKey, - ignored -> - new TornadoCausalAttentionPlan( - "attention-layer-" + layer, - executionBatchSize, - numHeads, - numKvHeads, - keyLength, - valueLength, - maxSequenceLength, - slidingWindow)); - long started = System.nanoTime(); - plan.execute(query, key, value, output, batchSize, startPosition); - totalNanos += System.nanoTime() - started; - calls++; - nextPositions.put(layer, startPosition + batchSize); - } - - public synchronized int planCount() { - return plans.size(); - } - - public synchronized long calls() { - return calls; - } - - public synchronized double totalMillis() { - return totalNanos / 1_000_000.0; - } - - @Override - public synchronized void rewind(int checkpoint) { - nextPositions.replaceAll((ignored, nextPosition) -> Math.min(nextPosition, checkpoint)); - } - - @Override - public synchronized void reset() { - nextPositions.clear(); - } - - @Override - public synchronized void close() { - if (closed) { - return; - } - RuntimeException failure = null; - for (TornadoCausalAttentionPlan plan : plans.values()) { - try { - plan.close(); - } catch (RuntimeException exception) { - if (failure == null) { - failure = exception; - } else { - failure.addSuppressed(exception); - } - } - } - plans.clear(); - nextPositions.clear(); - closed = true; - if (failure != null) { - throw failure; - } - } - - private record AttentionKey( - int layer, - int numHeads, - int numKvHeads, - int keyLength, - int valueLength, - int maxSequenceLength, - int slidingWindow) {} -} diff --git a/models-accelerator-bench/src/main/java/com/integrallis/models/accelerator/TornadoCausalAttentionPlan.java b/models-accelerator-bench/src/main/java/com/integrallis/models/accelerator/TornadoCausalAttentionPlan.java deleted file mode 100644 index 11b891e3..00000000 --- a/models-accelerator-bench/src/main/java/com/integrallis/models/accelerator/TornadoCausalAttentionPlan.java +++ /dev/null @@ -1,197 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.accelerator; - -import java.lang.foreign.MemorySegment; -import java.util.Arrays; -import uk.ac.manchester.tornado.api.GridScheduler; -import uk.ac.manchester.tornado.api.KernelContext; -import uk.ac.manchester.tornado.api.TaskGraph; -import uk.ac.manchester.tornado.api.TornadoExecutionPlan; -import uk.ac.manchester.tornado.api.WorkerGrid; -import uk.ac.manchester.tornado.api.WorkerGrid1D; -import uk.ac.manchester.tornado.api.enums.DataTransferMode; -import uk.ac.manchester.tornado.api.types.arrays.FloatArray; -import uk.ac.manchester.tornado.api.types.arrays.IntArray; - -/** One fixed-shape causal-attention plan with a persistent device KV cache. */ -final class TornadoCausalAttentionPlan implements AutoCloseable { - private final int executionBatchSize; - private final int queryDim; - private final int keyDim; - private final int valueDim; - private final int outputDim; - private final int maxSequenceLength; - private final float[] paddedQuery; - private final float[] paddedKey; - private final float[] paddedValue; - private final IntArray state; - private final FloatArray deviceQuery; - private final FloatArray deviceKey; - private final FloatArray deviceValue; - private final FloatArray deviceOutput; - private final TornadoExecutionPlan plan; - private long calls; - private long totalNanos; - - TornadoCausalAttentionPlan( - String name, - int executionBatchSize, - int numHeads, - int numKvHeads, - int keyLength, - int valueLength, - int maxSequenceLength, - int slidingWindow) { - if (executionBatchSize < 1 - || numHeads < 1 - || numKvHeads < 1 - || numHeads % numKvHeads != 0 - || keyLength < 1 - || valueLength != keyLength - || maxSequenceLength < executionBatchSize - || slidingWindow != 0) { - throw new IllegalArgumentException("invalid causal-attention shape"); - } - this.executionBatchSize = executionBatchSize; - this.queryDim = Math.multiplyExact(numHeads, keyLength); - this.keyDim = Math.multiplyExact(numKvHeads, keyLength); - this.valueDim = Math.multiplyExact(numKvHeads, valueLength); - this.outputDim = Math.multiplyExact(numHeads, valueLength); - this.maxSequenceLength = maxSequenceLength; - this.paddedQuery = new float[Math.multiplyExact(executionBatchSize, queryDim)]; - this.paddedKey = new float[Math.multiplyExact(executionBatchSize, keyDim)]; - this.paddedValue = new float[Math.multiplyExact(executionBatchSize, valueDim)]; - this.state = new IntArray(2); - this.deviceQuery = new FloatArray(paddedQuery.length); - this.deviceKey = new FloatArray(paddedKey.length); - this.deviceValue = new FloatArray(paddedValue.length); - FloatArray deviceKeyCache = new FloatArray(Math.multiplyExact(maxSequenceLength, keyDim)); - FloatArray deviceValueCache = new FloatArray(Math.multiplyExact(maxSequenceLength, valueDim)); - this.deviceOutput = new FloatArray(Math.multiplyExact(executionBatchSize, outputDim)); - KernelContext context = new KernelContext(); - TaskGraph graph = - new TaskGraph(name) - .transferToDevice( - DataTransferMode.FIRST_EXECUTION, context, deviceKeyCache, deviceValueCache) - .transferToDevice( - DataTransferMode.EVERY_EXECUTION, state, deviceQuery, deviceKey, deviceValue) - .task( - "store-kv", - CausalAttentionKernel::store, - state, - deviceKey, - deviceValue, - deviceKeyCache, - deviceValueCache, - keyDim, - valueDim) - .task( - "attention", - CausalAttentionKernel::attendTiled, - context, - state, - deviceQuery, - deviceKeyCache, - deviceValueCache, - deviceOutput, - numHeads, - keyLength, - keyDim, - numHeads / numKvHeads, - queryDim) - .transferToHost(DataTransferMode.EVERY_EXECUTION, deviceOutput); - int localSize = optimalLocalSize(keyLength); - int globalSize = - Math.multiplyExact(Math.multiplyExact(executionBatchSize, numHeads), localSize); - WorkerGrid workerGrid = new WorkerGrid1D(globalSize); - workerGrid.setGlobalWork(globalSize, 1, 1); - workerGrid.setLocalWork(localSize, 1, 1); - GridScheduler scheduler = new GridScheduler(name + ".attention", workerGrid); - this.plan = new TornadoExecutionPlan(graph.snapshot()).withGridScheduler(scheduler); - } - - void execute( - float[] query, - float[] key, - float[] value, - float[] output, - int actualBatchSize, - int startPosition) { - if (actualBatchSize < 1 || actualBatchSize > executionBatchSize) { - throw new IllegalArgumentException("actualBatchSize exceeds execution shape"); - } - if (startPosition < 0 || startPosition > maxSequenceLength - actualBatchSize) { - throw new IllegalArgumentException("prompt chunk exceeds attention context"); - } - stage(query, paddedQuery, actualBatchSize, queryDim, "query"); - stage(key, paddedKey, actualBatchSize, keyDim, "key"); - stage(value, paddedValue, actualBatchSize, valueDim, "value"); - if (output.length < Math.multiplyExact(actualBatchSize, outputDim)) { - throw new IllegalArgumentException("output storage does not match attention shape"); - } - deviceQuery.getSegment().copyFrom(MemorySegment.ofArray(paddedQuery)); - deviceKey.getSegment().copyFrom(MemorySegment.ofArray(paddedKey)); - deviceValue.getSegment().copyFrom(MemorySegment.ofArray(paddedValue)); - state.set(0, startPosition); - state.set(1, actualBatchSize); - long started = System.nanoTime(); - plan.execute(); - long byteSize = Math.multiplyExact((long) actualBatchSize * outputDim, Float.BYTES); - MemorySegment.ofArray(output) - .asSlice(0, byteSize) - .copyFrom(deviceOutput.getSegment().asSlice(0, byteSize)); - totalNanos += System.nanoTime() - started; - calls++; - } - - long calls() { - return calls; - } - - double totalMillis() { - return totalNanos / 1_000_000.0; - } - - private static void stage( - float[] source, float[] padded, int actualBatchSize, int dimension, String name) { - int actualEntries = Math.multiplyExact(actualBatchSize, dimension); - if (source.length < actualEntries) { - throw new IllegalArgumentException(name + " storage does not match attention shape"); - } - System.arraycopy(source, 0, padded, 0, actualEntries); - Arrays.fill(padded, actualEntries, padded.length, 0.0f); - } - - private static int optimalLocalSize(int dimension) { - int maximum = Math.min(dimension, 64); - for (int candidate = maximum; candidate >= 1; candidate--) { - if (dimension % candidate == 0) { - return candidate; - } - } - return 1; - } - - @Override - public void close() { - try { - plan.close(); - } catch (Exception exception) { - throw new IllegalStateException("could not close Tornado attention plan", exception); - } - } -} diff --git a/models-accelerator-bench/src/test/java/com/integrallis/models/accelerator/CausalAttentionKernelTest.java b/models-accelerator-bench/src/test/java/com/integrallis/models/accelerator/CausalAttentionKernelTest.java deleted file mode 100644 index 567ee9f1..00000000 --- a/models-accelerator-bench/src/test/java/com/integrallis/models/accelerator/CausalAttentionKernelTest.java +++ /dev/null @@ -1,181 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.accelerator; - -import static org.assertj.core.api.Assertions.assertThat; - -import java.util.Random; -import org.junit.jupiter.api.Test; -import uk.ac.manchester.tornado.api.types.arrays.FloatArray; -import uk.ac.manchester.tornado.api.types.arrays.IntArray; - -class CausalAttentionKernelTest { - private static final int EXECUTION_BATCH = 4; - private static final int HEADS = 4; - private static final int KV_HEADS = 2; - private static final int HEAD_LENGTH = 8; - private static final int QUERY_DIM = HEADS * HEAD_LENGTH; - private static final int KV_DIM = KV_HEADS * HEAD_LENGTH; - private static final int MAX_SEQUENCE = 12; - - @Test - void matchesGroupedQueryAttentionAcrossIncrementalPromptChunks() { - FloatArray keyCache = new FloatArray(MAX_SEQUENCE * KV_DIM); - FloatArray valueCache = new FloatArray(MAX_SEQUENCE * KV_DIM); - FloatArray scores = new FloatArray(EXECUTION_BATCH * HEADS * MAX_SEQUENCE); - FloatArray output = new FloatArray(EXECUTION_BATCH * QUERY_DIM); - IntArray state = new IntArray(2); - float[] referenceKeys = new float[MAX_SEQUENCE * KV_DIM]; - float[] referenceValues = new float[MAX_SEQUENCE * KV_DIM]; - - runAndCompareChunk( - 0, 4, 41L, keyCache, valueCache, scores, output, state, referenceKeys, referenceValues); - runAndCompareChunk( - 4, 3, 43L, keyCache, valueCache, scores, output, state, referenceKeys, referenceValues); - } - - private static void runAndCompareChunk( - int startPosition, - int actualBatch, - long seed, - FloatArray keyCache, - FloatArray valueCache, - FloatArray scores, - FloatArray output, - IntArray state, - float[] referenceKeys, - float[] referenceValues) { - float[] query = randomFloats(EXECUTION_BATCH * QUERY_DIM, seed); - float[] key = randomFloats(EXECUTION_BATCH * KV_DIM, seed + 1); - float[] value = randomFloats(EXECUTION_BATCH * KV_DIM, seed + 2); - state.set(0, startPosition); - state.set(1, actualBatch); - FloatArray deviceQuery = FloatArray.fromArray(query); - FloatArray deviceKey = FloatArray.fromArray(key); - FloatArray deviceValue = FloatArray.fromArray(value); - - CausalAttentionKernel.store( - state, deviceKey, deviceValue, keyCache, valueCache, KV_DIM, KV_DIM); - CausalAttentionKernel.attend( - state, - deviceQuery, - keyCache, - valueCache, - scores, - output, - HEADS, - KV_HEADS, - HEAD_LENGTH, - HEAD_LENGTH, - KV_DIM, - KV_DIM, - MAX_SEQUENCE, - 0); - - storeReference( - startPosition, actualBatch, key, value, referenceKeys, referenceValues, KV_DIM, KV_DIM); - float[] expected = - referenceAttention( - startPosition, - actualBatch, - query, - referenceKeys, - referenceValues, - HEADS, - KV_HEADS, - HEAD_LENGTH, - HEAD_LENGTH, - KV_DIM, - KV_DIM); - assertThat(output.toHeapArray()) - .usingComparatorWithPrecision(1.0e-5f) - .containsExactly(expected); - } - - private static float[] randomFloats(int length, long seed) { - Random random = new Random(seed); - float[] values = new float[length]; - for (int index = 0; index < values.length; index++) { - values[index] = random.nextFloat(-1.0f, 1.0f); - } - return values; - } - - private static void storeReference( - int startPosition, - int actualBatch, - float[] key, - float[] value, - float[] keyCache, - float[] valueCache, - int keyDim, - int valueDim) { - for (int batch = 0; batch < actualBatch; batch++) { - System.arraycopy(key, batch * keyDim, keyCache, (startPosition + batch) * keyDim, keyDim); - System.arraycopy( - value, batch * valueDim, valueCache, (startPosition + batch) * valueDim, valueDim); - } - } - - private static float[] referenceAttention( - int startPosition, - int actualBatch, - float[] query, - float[] keyCache, - float[] valueCache, - int heads, - int kvHeads, - int keyLength, - int valueLength, - int keyDim, - int valueDim) { - float[] output = new float[EXECUTION_BATCH * heads * valueLength]; - float scale = (float) (1.0 / Math.sqrt(keyLength)); - int groupSize = heads / kvHeads; - for (int batch = 0; batch < actualBatch; batch++) { - int position = startPosition + batch; - for (int head = 0; head < heads; head++) { - int kvHead = head / groupSize; - float[] attention = new float[position + 1]; - float max = Float.NEGATIVE_INFINITY; - for (int cached = 0; cached <= position; cached++) { - float dot = 0.0f; - for (int column = 0; column < keyLength; column++) { - dot += - query[batch * heads * keyLength + head * keyLength + column] - * keyCache[cached * keyDim + kvHead * keyLength + column]; - } - attention[cached] = dot * scale; - max = Math.max(max, attention[cached]); - } - float sum = 0.0f; - for (int cached = 0; cached <= position; cached++) { - attention[cached] = (float) Math.exp(attention[cached] - max); - sum += attention[cached]; - } - for (int column = 0; column < valueLength; column++) { - float weighted = 0.0f; - for (int cached = 0; cached <= position; cached++) { - weighted += - attention[cached] * valueCache[cached * valueDim + kvHead * valueLength + column]; - } - output[batch * heads * valueLength + head * valueLength + column] = weighted / sum; - } - } - } - return output; - } -} diff --git a/models-accelerator-bench/src/test/java/com/integrallis/models/accelerator/Q8ProjectionKernelTest.java b/models-accelerator-bench/src/test/java/com/integrallis/models/accelerator/Q8ProjectionKernelTest.java deleted file mode 100644 index 4ea3721f..00000000 --- a/models-accelerator-bench/src/test/java/com/integrallis/models/accelerator/Q8ProjectionKernelTest.java +++ /dev/null @@ -1,100 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.accelerator; - -import static org.assertj.core.api.Assertions.assertThat; -import static org.assertj.core.api.Assertions.assertThatThrownBy; -import static org.assertj.core.api.Assertions.within; - -import com.integrallis.vectors.core.VectorUtil; -import java.lang.foreign.MemorySegment; -import java.util.Random; -import org.junit.jupiter.api.Test; -import uk.ac.manchester.tornado.api.types.arrays.ByteArray; -import uk.ac.manchester.tornado.api.types.arrays.FloatArray; - -class Q8ProjectionKernelTest { - - @Test - void matchesVectorApiDequantizedProjectionAcrossBatches() { - int batchSize = 3; - int rows = 17; - int cols = 96; - byte[] weights = randomQ8Matrix(rows, cols, 23L); - float[] input = randomFloats(batchSize * cols, 29L); - float[] expected = vectorApiProjection(weights, input, batchSize, rows, cols); - ByteArray deviceWeights = ByteArray.fromArray(weights); - FloatArray deviceInput = FloatArray.fromArray(input); - FloatArray deviceOutput = new FloatArray(batchSize * rows); - - Q8ProjectionKernel.multiply(deviceWeights, deviceInput, deviceOutput, batchSize, rows, cols); - - float[] actual = deviceOutput.toHeapArray(); - assertThat(actual).hasSameSizeAs(expected); - for (int index = 0; index < expected.length; index++) { - assertThat(actual[index]).isCloseTo(expected[index], within(2.0e-5f)); - } - } - - @Test - void rejectsColumnsThatAreNotQ8BlockAligned() { - assertThatThrownBy( - () -> - Q8ProjectionKernel.validate( - new ByteArray(34), new FloatArray(31), new FloatArray(1), 1, 1, 31)) - .isInstanceOf(IllegalArgumentException.class) - .hasMessageContaining("multiple of 32"); - } - - private static float[] vectorApiProjection( - byte[] weights, float[] input, int batchSize, int rows, int cols) { - float[] output = new float[batchSize * rows]; - MemorySegment weightSegment = MemorySegment.ofArray(weights); - for (int batch = 0; batch < batchSize; batch++) { - float[] query = new float[cols]; - float[] projected = new float[rows]; - System.arraycopy(input, batch * cols, query, 0, cols); - VectorUtil.ggufQ8_0BatchDotProduct(query, weightSegment, rows, cols, projected); - System.arraycopy(projected, 0, output, batch * rows, rows); - } - return output; - } - - private static byte[] randomQ8Matrix(int rows, int cols, long seed) { - Random random = new Random(seed); - int blocks = rows * cols / 32; - byte[] weights = new byte[blocks * 34]; - for (int block = 0; block < blocks; block++) { - short scale = Float.floatToFloat16(0.001f + random.nextFloat() * 0.05f); - int offset = block * 34; - weights[offset] = (byte) scale; - weights[offset + 1] = (byte) (scale >>> 8); - for (int quant = 0; quant < 32; quant++) { - weights[offset + 2 + quant] = (byte) (random.nextInt(255) - 127); - } - } - return weights; - } - - private static float[] randomFloats(int length, long seed) { - Random random = new Random(seed); - float[] values = new float[length]; - for (int index = 0; index < length; index++) { - values[index] = random.nextFloat(-2.0f, 2.0f); - } - return values; - } -} diff --git a/models-accelerator-bench/src/test/java/com/integrallis/models/accelerator/QwenFullModelExperimentTest.java b/models-accelerator-bench/src/test/java/com/integrallis/models/accelerator/QwenFullModelExperimentTest.java deleted file mode 100644 index d624f0cf..00000000 --- a/models-accelerator-bench/src/test/java/com/integrallis/models/accelerator/QwenFullModelExperimentTest.java +++ /dev/null @@ -1,37 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.accelerator; - -import static org.assertj.core.api.Assertions.assertThat; -import static org.assertj.core.api.Assertions.assertThatThrownBy; - -import org.junit.jupiter.api.Test; - -class QwenFullModelExperimentTest { - - @Test - void buildsAFullFixedShapeReadinessPromptFromRealTokens() { - assertThat(QwenFullModelExperiment.readinessTokens(new int[] {7, 8, 9}, 8)) - .containsExactly(7, 8, 9, 7, 8, 9, 7, 8); - } - - @Test - void refusesToPrepareWithoutAnyRealTokens() { - assertThatThrownBy(() -> QwenFullModelExperiment.readinessTokens(new int[0], 8)) - .isInstanceOf(IllegalArgumentException.class) - .hasMessageContaining("tokens"); - } -} diff --git a/models-bench/README.md b/models-bench/README.md index 47209f5c..43adac99 100644 --- a/models-bench/README.md +++ b/models-bench/README.md @@ -637,40 +637,17 @@ The requested context must hold the prompt plus the larger of the warmup and mea The resolved artifact is validated before loading, and the command rejects token IDs outside the model vocabulary. -## Accelerator profile (GPU gate) +## Accelerator profile (GPU gate) — removed -Run prefill and greedy decode for a GGUF file on the TornadoVM backend and record the evidence a -hardware gate needs. Unlike the other profiles this one takes a GGUF path directly rather than a -catalog model id, because the accelerator is selected from the file's size and the device inventory. +`backend-tornado` and this command were removed in 0.3.53. Models ships one GPU implementation, +`backend-cuda`, because on the same RTX 4090 and the same model it measured exact token parity +(1280 of 1280 token ids identical) and 31.26 tok/s decode against a 5.06 tok/s CPU control, while +the TornadoVM arm managed 6.49 tok/s over 36 failed `cuLaunchKernel` calls that its report did not +record. The evidence for both is in `benchmark-results/2026-10-07-g1-parity`, +`benchmark-results/2026-10-07-g4-dualpath` and `benchmark-results/2026-10-07-tornado-removal`. -```shell -./gradlew :models-bench:installDist - -lib=models-bench/build/install/models-bench/lib -classpath=$(printf '%s:' "$lib"/*.jar) - -tornado -cp "$classpath" \ - --params="accelerator-profile --model ~/.jvllm/models/Qwen3-27B-Q4_K_M.gguf \ - --tokens 64 --batch 32 --require true \ - --output build/reports/inference/accelerator-profile.json" \ - com.integrallis.models.bench.InferenceBenchmarkCli -``` - -The command reports, to stdout and as JSON, whether the accelerator was selected, the device name, -the fallback reason when it was not, the eager-readiness time, the number of retained device -execution plans, prefill and decode tokens per second, and **how many projections reached the device -per GGUF weight format**. That last figure is the point of the command: a mixed-format model such as -a `Q4_K_M` build contains both Q4_K and Q6_K tensors, and a run whose `Q6_K` count is zero routed -only part of the model while still producing a plausible throughput number. - -`--require true` turns an unavailable or ineligible accelerator into a hard failure. Without it, and -without a qualified device, the command still completes and the report records -`accelerated=false` with the reason — which is the useful CPU control arm. - -Options: `--prompt` / `--prompt-file`, `--tokens` (generated tokens, default 64), `--warmups` -(default 1), `--context` (default 2048), `--batch` (fixed device prefill batch, default 32), -`--decode` (accelerate single-token decode, default true), `--eager` (compile plans before the first -visible request, default true), `--require` (default false), `--output`. +For the device gates use `cuda-kernel-gate` — `--mode capability`, `--mode parity` (G1) and +`--mode decode` (G4). It needs no external runtime installation. ## In-process comparison run diff --git a/models-bench/build.gradle.kts b/models-bench/build.gradle.kts index 80d637cb..76fd2909 100644 --- a/models-bench/build.gradle.kts +++ b/models-bench/build.gradle.kts @@ -67,11 +67,6 @@ dependencies { implementation(project(":models-runtime")) implementation(project(":backend-java")) implementation(project(":models-router")) - // The accelerator profile gate drives the Tornado backend directly. backend-tornado keeps - // tornado-api off its public API, so nothing here needs it at compile time; the TornadoVM - // launcher supplies the device runtime on a GPU host. Without one, TornadoBackend.open catches - // the LinkageError and reports a Vector API fallback, which is what the report should say. - implementation(project(":backend-tornado")) if (nativeBenchmarkRuntime && !aggregateNativeRelease) { // The platform artifact contains only the compiled library and metadata. The Java FFM // bridge lives in backend-native's ordinary JAR and must be present as well. diff --git a/models-bench/src/main/java/com/integrallis/models/bench/AcceleratorProfileCli.java b/models-bench/src/main/java/com/integrallis/models/bench/AcceleratorProfileCli.java deleted file mode 100644 index bacc8450..00000000 --- a/models-bench/src/main/java/com/integrallis/models/bench/AcceleratorProfileCli.java +++ /dev/null @@ -1,265 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.bench; - -import com.fasterxml.jackson.databind.ObjectMapper; -import com.fasterxml.jackson.databind.SerializationFeature; -import com.integrallis.models.api.Tokenizer; -import com.integrallis.models.backend.purejava.PureJavaBackend; -import com.integrallis.models.backend.tornado.TornadoBackend; -import com.integrallis.models.backend.tornado.TornadoBackendOptions; -import com.integrallis.models.backend.tornado.TornadoBackendRuntime; -import com.integrallis.models.backend.tornado.TornadoBackendStatus; -import java.io.IOException; -import java.nio.file.Files; -import java.nio.file.Path; -import java.time.Instant; -import java.util.LinkedHashMap; -import java.util.Map; -import java.util.Set; - -/** - * Runs one prefill and one greedy decode of a GGUF model through the TornadoVM backend and reports - * the evidence a hardware gate needs. - * - *

The report is deliberately shaped like the other gates: it records the selected device and the - * fallback reason if any, the readiness cost the eager plan compilation paid, prefill and decode - * throughput, and — the point of this command — the number of projections that actually reached the - * device per GGUF weight format. A Q4_K_M model whose {@code Q6_K} count is zero routed only - * its Q4_K tensors, which looks like a working accelerator on the throughput line alone. - * - *

Off a qualified GPU this command still runs: the Tornado backend falls back to the Vector API - * and the report says so, with {@code accelerated=false} and the reason. Pass {@code --require - * true} to make an unavailable accelerator a hard failure instead. - */ -final class AcceleratorProfileCli { - - private static final Set OPTIONS = - Set.of( - "model", - "prompt", - "prompt-file", - "tokens", - "warmups", - "context", - "batch", - "decode", - "eager", - "require", - "output"); - private static final String DEFAULT_PROMPT = - "Explain why a GPU projection kernel must be checked against the CPU kernel it replaces " - + "before any throughput number from it is worth reporting."; - - private AcceleratorProfileCli() {} - - static void run(String[] args) throws IOException { - Configuration configuration = parse(args); - System.setProperty( - PureJavaBackend.MAX_CONTEXT_LENGTH_PROPERTY, - Integer.toString(configuration.contextLength())); - TornadoBackendOptions options = - new TornadoBackendOptions( - configuration.decode(), - configuration.eager(), - configuration.require(), - configuration.batch()); - try (TornadoBackendRuntime runtime = TornadoBackend.open(configuration.model(), options)) { - Report report = profile(runtime, configuration); - write(report, configuration.output()); - print(report, configuration.output()); - } - } - - static Configuration parse(String[] args) throws IOException { - Map values = BenchmarkCliArguments.parse(args, OPTIONS); - String model = values.get("model"); - if (model == null || model.isBlank()) { - throw new IllegalArgumentException("--model must be a GGUF file path"); - } - Path modelPath = Path.of(model).toAbsolutePath().normalize(); - if (!Files.isRegularFile(modelPath)) { - throw new IllegalArgumentException("--model must be a regular file: " + modelPath); - } - String prompt = - values.containsKey("prompt-file") - ? Files.readString(Path.of(values.get("prompt-file"))) - : values.getOrDefault("prompt", DEFAULT_PROMPT); - return new Configuration( - modelPath, - prompt, - BenchmarkCliArguments.integer(values, "tokens", 64), - BenchmarkCliArguments.integer(values, "warmups", 1), - BenchmarkCliArguments.integer(values, "context", 2_048), - BenchmarkCliArguments.integer(values, "batch", 32), - flag(values, "decode", true), - flag(values, "eager", true), - flag(values, "require", false), - Path.of(values.getOrDefault("output", "build/reports/inference/accelerator-profile.json"))); - } - - static Report profile(TornadoBackendRuntime runtime, Configuration configuration) { - PureJavaBackend backend = runtime.backend(); - Tokenizer tokenizer = backend.tokenizer(); - int[] promptTokens = tokenizer.encode(configuration.prompt()); - if (promptTokens.length == 0) { - throw new IllegalArgumentException("prompt produced no tokens"); - } - if (promptTokens.length + configuration.tokens() > configuration.contextLength()) { - throw new IllegalArgumentException( - "context length " - + configuration.contextLength() - + " cannot hold " - + promptTokens.length - + " prompt tokens plus " - + configuration.tokens() - + " generated tokens"); - } - - for (int warmup = 0; warmup < configuration.warmups(); warmup++) { - backend.reset(); - float[] logits = backend.prefill(promptTokens, 0); - generate(backend, logits, promptTokens.length, Math.min(4, configuration.tokens())); - } - - backend.reset(); - long prefillStart = System.nanoTime(); - float[] logits = backend.prefill(promptTokens, 0); - long prefillNanos = System.nanoTime() - prefillStart; - - long decodeStart = System.nanoTime(); - int generated = generate(backend, logits, promptTokens.length, configuration.tokens()); - long decodeNanos = System.nanoTime() - decodeStart; - - TornadoBackendStatus status = runtime.status(); - return new Report( - Instant.now().toString(), - configuration.model().toString(), - status.accelerated(), - status.device(), - status.reason(), - status.requiredDeviceBytes(), - status.readinessTime().toMillis(), - runtime.projectionPlanCount(), - new LinkedHashMap<>(runtime.routedProjectionsByFormat()), - promptTokens.length, - generated, - perSecond(promptTokens.length, prefillNanos), - perSecond(generated, decodeNanos), - prefillNanos / 1_000_000.0, - decodeNanos / 1_000_000.0); - } - - private static int generate(PureJavaBackend backend, float[] logits, int position, int tokens) { - float[] current = logits; - int generated = 0; - for (int index = 0; index < tokens; index++) { - int next = argmax(current); - generated++; - if (index + 1 < tokens) { - current = backend.forward(next, position + index); - } - } - return generated; - } - - private static int argmax(float[] logits) { - int best = 0; - for (int index = 1; index < logits.length; index++) { - if (logits[index] > logits[best]) { - best = index; - } - } - return best; - } - - private static double perSecond(int count, long nanos) { - return nanos <= 0 ? 0.0 : count * 1_000_000_000.0 / nanos; - } - - private static void write(Report report, Path output) throws IOException { - Path parent = output.toAbsolutePath().getParent(); - if (parent != null) { - Files.createDirectories(parent); - } - new ObjectMapper() - .enable(SerializationFeature.INDENT_OUTPUT) - .writeValue(output.toFile(), report); - } - - private static void print(Report report, Path output) { - System.out.printf( - "accelerator profile: accelerated=%s device=%s reason=%s readiness=%d ms plans=%d%n" - + " prompt=%d tokens prefill=%.2f tok/s (%.1f ms) " - + "decode=%d tokens %.2f tok/s (%.1f ms)%n" - + " routed projections by format: %s%nreport: %s%n", - report.accelerated(), - report.device(), - report.reason(), - report.readinessMillis(), - report.projectionPlans(), - report.promptTokens(), - report.prefillTokensPerSecond(), - report.prefillMillis(), - report.generatedTokens(), - report.decodeTokensPerSecond(), - report.decodeMillis(), - report.routedProjectionsByFormat().isEmpty() - ? "none (nothing reached the device)" - : report.routedProjectionsByFormat(), - output.toAbsolutePath()); - } - - private static boolean flag(Map values, String name, boolean defaultValue) { - String value = values.get(name); - if (value == null) { - return defaultValue; - } - if ("true".equalsIgnoreCase(value) || "false".equalsIgnoreCase(value)) { - return Boolean.parseBoolean(value); - } - throw new IllegalArgumentException("--" + name + " must be true or false: " + value); - } - - record Configuration( - Path model, - String prompt, - int tokens, - int warmups, - int contextLength, - int batch, - boolean decode, - boolean eager, - boolean require, - Path output) {} - - record Report( - String timestamp, - String model, - boolean accelerated, - String device, - String reason, - long requiredBytes, - long readinessMillis, - int projectionPlans, - Map routedProjectionsByFormat, - int promptTokens, - int generatedTokens, - double prefillTokensPerSecond, - double decodeTokensPerSecond, - double prefillMillis, - double decodeMillis) {} -} diff --git a/models-bench/src/main/java/com/integrallis/models/bench/InferenceBenchmarkCli.java b/models-bench/src/main/java/com/integrallis/models/bench/InferenceBenchmarkCli.java index 2fe04d1c..bde4bbbe 100644 --- a/models-bench/src/main/java/com/integrallis/models/bench/InferenceBenchmarkCli.java +++ b/models-bench/src/main/java/com/integrallis/models/bench/InferenceBenchmarkCli.java @@ -91,10 +91,6 @@ public static void main(String[] args) throws Exception { RaggedPrefillProfileCli.run(Arrays.copyOfRange(args, 1, args.length)); return; } - if (args.length > 0 && "accelerator-profile".equals(args[0])) { - AcceleratorProfileCli.run(Arrays.copyOfRange(args, 1, args.length)); - return; - } if (args.length > 0 && "compare".equals(args[0])) { BenchmarkComparisonCli.run(Arrays.copyOfRange(args, 1, args.length)); return; diff --git a/models-bench/src/test/java/com/integrallis/models/bench/AcceleratorProfileCliTest.java b/models-bench/src/test/java/com/integrallis/models/bench/AcceleratorProfileCliTest.java deleted file mode 100644 index 482f6a81..00000000 --- a/models-bench/src/test/java/com/integrallis/models/bench/AcceleratorProfileCliTest.java +++ /dev/null @@ -1,101 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.bench; - -import static org.assertj.core.api.Assertions.assertThat; -import static org.assertj.core.api.Assertions.assertThatThrownBy; - -import java.io.IOException; -import java.nio.file.Files; -import java.nio.file.Path; -import org.junit.jupiter.api.Test; -import org.junit.jupiter.api.io.TempDir; - -class AcceleratorProfileCliTest { - - @TempDir Path directory; - - @Test - void parsesTheAcceleratorGateOptions() throws IOException { - Path model = Files.writeString(directory.resolve("model.gguf"), "not a real model"); - - AcceleratorProfileCli.Configuration configuration = - AcceleratorProfileCli.parse( - new String[] { - "--model", model.toString(), - "--prompt", "measure it", - "--tokens", "16", - "--warmups", "0", - "--context", "512", - "--batch", "8", - "--decode", "false", - "--eager", "false", - "--require", "true", - "--output", directory.resolve("gate.json").toString() - }); - - assertThat(configuration.model()).isEqualTo(model.toAbsolutePath().normalize()); - assertThat(configuration.prompt()).isEqualTo("measure it"); - assertThat(configuration.tokens()).isEqualTo(16); - assertThat(configuration.warmups()).isZero(); - assertThat(configuration.contextLength()).isEqualTo(512); - assertThat(configuration.batch()).isEqualTo(8); - assertThat(configuration.decode()).isFalse(); - assertThat(configuration.eager()).isFalse(); - assertThat(configuration.require()).isTrue(); - assertThat(configuration.output()).hasFileName("gate.json"); - } - - @Test - void defaultsToAnAcceleratedPrefillAndDecodeGate() throws IOException { - Path model = Files.writeString(directory.resolve("model.gguf"), "not a real model"); - - AcceleratorProfileCli.Configuration configuration = - AcceleratorProfileCli.parse(new String[] {"--model", model.toString()}); - - assertThat(configuration.decode()).isTrue(); - assertThat(configuration.eager()).isTrue(); - assertThat(configuration.require()).isFalse(); - assertThat(configuration.batch()).isEqualTo(32); - assertThat(configuration.tokens()).isEqualTo(64); - assertThat(configuration.prompt()).isNotBlank(); - } - - @Test - void rejectsAModelPathThatIsNotAFile() { - assertThatThrownBy( - () -> - AcceleratorProfileCli.parse( - new String[] {"--model", directory.resolve("missing.gguf").toString()})) - .isInstanceOf(IllegalArgumentException.class) - .hasMessageContaining("regular file"); - assertThatThrownBy(() -> AcceleratorProfileCli.parse(new String[0])) - .isInstanceOf(IllegalArgumentException.class) - .hasMessageContaining("--model"); - } - - @Test - void rejectsNonBooleanFlagValues() throws IOException { - Path model = Files.writeString(directory.resolve("model.gguf"), "not a real model"); - - assertThatThrownBy( - () -> - AcceleratorProfileCli.parse( - new String[] {"--model", model.toString(), "--require", "yes"})) - .isInstanceOf(IllegalArgumentException.class) - .hasMessageContaining("--require must be true or false"); - } -} diff --git a/settings.gradle.kts b/settings.gradle.kts index 7199a151..dd271e80 100644 --- a/settings.gradle.kts +++ b/settings.gradle.kts @@ -18,7 +18,6 @@ include("models-decisions") // --- Backends --- include("backend-java") -include("backend-tornado") include("models-backend-onnx") include("backend-native") include("backend-cuda") From 3cd08285f25f82f1c5d1648c8f6a3015d62af839 Mon Sep 17 00:00:00 2001 From: Brian Sam-Bodden Date: Wed, 7 Oct 2026 15:02:52 -0700 Subject: [PATCH 5/5] Bump every release-metadata surface to 0.3.53, not just gradle.properties CI failed on :verifyReleaseMetadata with "Antora display_version must match 0.3.53". The version bump had changed gradle.properties alone, and that task exists because a release where the published coordinate and the user-facing version disagree is a release that documents the wrong thing. Running it locally then surfaced two more classes it checks, in order: Antora display_version docs/content/antora.yml models-version attribute docs/content/antora.yml docs package version docs/package.json, docs/package-lock.json landing page pill and snippet docs/landing/index.html native crate version model-kernels Cargo.toml and Cargo.lock published coordinates in prose models-backend-apple/README.md, models-rag/README.md notebook release defaults notebooks/.env.example, notebooks/docker-compose.yml, notebooks/jupyter/prepare-classpath.sh, notebooks/README.md The Cargo.lock edit is anchored on the jmodels-kernels package entry rather than applied globally, so a dependency that happens to carry the same version string is not rewritten. The cause of the miss is worth recording: the local gate was spotlessCheck build -x :docs:build, and verifyReleaseMetadata hangs off complianceCheck, which the release workflow runs and that command does not. complianceCheck now passes locally, which is the check that should have been run before pushing a version bump. The merge-and-release chain stopped on the failing check instead of merging, which is what it was gated to do. --- backend-native/src/main/rust/model-kernels/Cargo.lock | 2 +- backend-native/src/main/rust/model-kernels/Cargo.toml | 2 +- docs/content/antora.yml | 4 ++-- docs/landing/index.html | 4 ++-- docs/package-lock.json | 4 ++-- docs/package.json | 2 +- models-backend-apple/README.md | 2 +- models-rag/README.md | 2 +- notebooks/.env.example | 2 +- notebooks/README.md | 8 ++++---- notebooks/docker-compose.yml | 2 +- notebooks/jupyter/prepare-classpath.sh | 2 +- 12 files changed, 18 insertions(+), 18 deletions(-) diff --git a/backend-native/src/main/rust/model-kernels/Cargo.lock b/backend-native/src/main/rust/model-kernels/Cargo.lock index caf41977..117405e4 100644 --- a/backend-native/src/main/rust/model-kernels/Cargo.lock +++ b/backend-native/src/main/rust/model-kernels/Cargo.lock @@ -4,4 +4,4 @@ version = 4 [[package]] name = "jmodels-kernels" -version = "0.3.52" +version = "0.3.53" diff --git a/backend-native/src/main/rust/model-kernels/Cargo.toml b/backend-native/src/main/rust/model-kernels/Cargo.toml index 646fc003..94bb99d1 100644 --- a/backend-native/src/main/rust/model-kernels/Cargo.toml +++ b/backend-native/src/main/rust/model-kernels/Cargo.toml @@ -1,6 +1,6 @@ [package] name = "jmodels-kernels" -version = "0.3.52" +version = "0.3.53" edition = "2024" license = "Apache-2.0" publish = false diff --git a/docs/content/antora.yml b/docs/content/antora.yml index 4dcabe8f..b7787933 100644 --- a/docs/content/antora.yml +++ b/docs/content/antora.yml @@ -1,7 +1,7 @@ name: models title: Models version: 'current' -display_version: '0.3.52' +display_version: '0.3.53' prerelease: false start_page: ROOT:index.adoc nav: @@ -10,7 +10,7 @@ asciidoc: attributes: source-language: java source-highlighter: highlight.js - models-version: '0.3.52' + models-version: '0.3.53' modeljars-version: '0.1.31' vectors-version: '0.1.28' url-models-github: https://github.com/integrallis/models diff --git a/docs/landing/index.html b/docs/landing/index.html index b25aed44..933acd23 100644 --- a/docs/landing/index.html +++ b/docs/landing/index.html @@ -36,7 +36,7 @@ / models - v0.3.52 + v0.3.53