diff --git a/.gitignore b/.gitignore index 6853079b0..d2bc96110 100644 --- a/.gitignore +++ b/.gitignore @@ -50,3 +50,7 @@ docs/content/modules/ROOT/attachments/javadoc/ # Model cache directory **/.jvllm/ + +# Agent worktrees and session scratch. These are embedded git repositories; git add -A +# swept one into a release commit once. +.claude/ diff --git a/CHANGELOG.md b/CHANGELOG.md index b68e4d366..2f0d108dc 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -4,6 +4,75 @@ All notable changes to models are documented here. ## [Unreleased] +## [0.3.53] - 2026-10-07 + +### Removed + +- **`backend-tornado` is removed.** Models ships one GPU implementation, and this was the weaker of + two. Measured on the same RTX 4090 and the same model on the same day: `backend-cuda` reached + exact token parity (1280 of 1280 token ids identical) and **31.26 tok/s** decode against a + **5.06 tok/s** Vector API control, with 344 ms readiness and no device errors. The TornadoVM arm + managed **6.49 tok/s** -- barely above the CPU path it exists to accelerate -- over **36 failed + `cuLaunchKernel` calls** (`CUDA_ERROR_LAUNCH_OUT_OF_RESOURCES`) that its report did not record, + with 32,171 ms readiness. It also could not be self-contained: the Maven artifact cannot carry + TornadoVM's device runtime, so an application had to install a matching distribution and launch + through its own launcher, which defeats "add a dependency and get acceleration". Evidence: + `benchmark-results/2026-10-07-tornado-removal`. +- Removed with it, because they existed only to serve it: the `accelerator-profile` command of + `models-bench`, the TornadoVM benchmark arm in `models-accelerator-bench`, the `tornado-api` and + `tornado-runtime` dependencies, and the release precondition in `RELEASING.md` that required a + loader and parity gate on each NVIDIA profile under `models-accelerator-bench/results/`. That + precondition existed for `backend-tornado`; `backend-cuda`'s gates are + `cuda-kernel-gate --mode capability|parity|decode` and need no external runtime. +- The KV-ridge experiment in `models-accelerator-bench` is unaffected, and the August measurement + records that name `backend-tornado` are kept as written -- they are what was measured then. + +### Added + +- `backend-cuda` is now published. It ships one `sm_80` PTX module serving every device of compute + capability 8.0 or above, carried in the jar under `META-INF/models/cuda/` with a SHA-256 the Java + loader recomputes. **It is opt-in and nothing activates it by accident**: there is no + `META-INF/services` entry, so no `ServiceLoader` discovers it, a consumer has to call + `CudaGgufBatchedMatrixKernel.open()` and inject the kernel, and `-Dmodels.cuda.disabled=true` is a + kill switch on top of that. Having the jar on the classpath does not route anything to a device. +- `CudaRoutingCounters.declined(type, stage, reason, rows, cols)`, with the shape in the key, and + `declinedProjections` / `totalDeclinedProjections` in the gate report. A projection that answered + no to `isEligible` and left through the Java branch previously incremented nothing, so an empty + refusals map read as "everything ran on the device" when it only meant "nothing hit an explicit + refusal path". + +### Fixed + +- **The FFN gate and up projections never reached the device.** `CudaGgufBatchedMatrixKernel` did + not override `multiplyDual`, so `isDualEligible` inherited the SPI default of `false` and + `LlamaForwardPass.dualMatmulDispatch` sent both to `TensorOps.ggufDualMatmul` on every layer of + every token. On Granite 4.1 3B they are 8192x2560 each -- 53.3% of a layer's projection + arithmetic. The triple path for query, key and value was implemented; the dual path was not, and + nothing recorded the asymmetry. + +### Measured + +Both device gates pass, on an RTX 4090 at compute capability 8.9, Granite 4.1 3B Q4_K_M, 20 prompts: + +| gate | result | evidence | +| --- | --- | --- | +| G1 token parity | **passed** -- 1280 token ids identical | `benchmark-results/2026-10-07-g1-parity` | +| G4 decode speed | **passed** -- 6.184x against a 3.00x gate | `benchmark-results/2026-10-07-g4-dualpath` | + +G4 moved from 1.788x to 6.184x, decode from 9.02 to 31.26 tok/s, with a control arm that moved +0.2% and a byte-identical PTX module either side. Routing the two projections raised dispatch -- +launches 241 to 321, transfers 402 to 522 per decode step -- and it did not matter. The earlier +reading of 1.788x as a dispatch ceiling was wrong; the binding constraint was the unimplemented +path. + +**One host and one model.** G1 has no tolerance, so each qualifying hardware profile and each +architecture family earns its own run before the claim generalises. + +The published surface changes in two ways and no others: `backend-cuda` is added and +`backend-tornado` is removed. No remaining published module's Java behaviour changed -- the rest of +`v0.3.52..HEAD` touches only benchmark applications and evidence. + + ## [0.3.52] - 2026-10-06 ### Fixed diff --git a/README.md b/README.md index a925e1880..e97f93249 100644 --- a/README.md +++ b/README.md @@ -34,9 +34,10 @@ tokens. Models implements that pipeline on Java 25 and uses the Vector API for CPU SIMD execution: - `backend-java` executes every inference kernel in Java. -- `backend-tornado` optionally compiles the Java Q4_0 projection kernels for a - qualified NVIDIA GPU. It keeps the Models graph in-process and falls back to - the Vector API when the device or artifact is not eligible. +- `backend-cuda` optionally runs the K-quant projections and grouped-query decode + attention on a qualified NVIDIA GPU, through our own Rust kernels compiled to + PTX. It keeps the Models graph in-process and falls back to the Vector API when + the device or artifact is not eligible. - `backend-native` runs the same Java 25 and Vector API pipeline, substituting only selected, measured bottleneck kernels with a small Models-owned Rust library through Java's Foreign Function and Memory (FFM) API. @@ -212,9 +213,11 @@ dependencies { } ``` -For qualified NVIDIA acceleration, add `backend-tornado` and launch with a -matching TornadoVM PTX runtime. The default loader performs eager readiness and -uses the Vector API when the GPU cannot safely retain the compiled plans. See +For qualified NVIDIA acceleration, add `backend-cuda`. One `sm_80` PTX module +ships inside the jar for every device of compute capability 8.0 or above, so no +external runtime is installed; the path is opt-in through +`CudaGgufBatchedMatrixKernel.open()` and falls back to the Vector API when the +device or artifact is not eligible. See [Java GPU acceleration](https://integrallis.github.io/models/docs/models/current/gpu-acceleration.html). Use Apple's on-device system model on a supported Apple Silicon Mac: @@ -375,7 +378,7 @@ documented in [Execution planning](https://integrallis.github.io/models/docs/mod | Model routing | `models-router` | adaptive selection and failover across in-process and hosted clients, with hard per-request capability and data-boundary requirements, without a provider SDK dependency | | Vector storage | `models-embedding` | optional bridge to `vectors` | | Apple on-device model | `backend-apple` | Apple Foundation Models through Java FFM | -| Java GPU acceleration | `backend-tornado` | optional Java-authored Q4_0 projections on qualified NVIDIA GPUs | +| NVIDIA GPU acceleration | `backend-cuda` | optional Rust-authored PTX K-quant projections and decode attention | These adapters are implemented and tested against the same backend contracts; they do not select hidden inference paths. Their framework dependencies are @@ -409,7 +412,7 @@ RAG, Javadocs, and release testing. - [Executable Java notebooks](notebooks/README.md) - [Apple Foundation Models bridge](models-backend-apple/README.md) -- [Java GPU acceleration](backend-tornado/README.md) +- [NVIDIA GPU acceleration](backend-cuda/README.md) - [Native kernel backend](backend-native/README.md) ## Build diff --git a/RELEASING.md b/RELEASING.md index c87f34255..990877da1 100644 --- a/RELEASING.md +++ b/RELEASING.md @@ -5,12 +5,25 @@ JReleaser signs and validates one Maven Central bundle, and the workflow creates the GitHub release. The publication allowlist contains `models-api`, `models-runtime`, `models`, -`models-rag`, `models-semantic-order`, `backend-java`, `backend-tornado`, `backend-native`, -`backend-apple`, `models-langchain4j`, `models-spring-ai`, +`models-rag`, `models-semantic-order`, `backend-java`, `backend-native`, +`backend-cuda`, `backend-apple`, `models-langchain4j`, `models-spring-ai`, `models-spring-boot-starter`, `models-embedding`, `models-audio`, `models-router`, and `models-decisions`. Benchmark applications, documentation tooling, and modules containing only package scaffolding are not published. +`backend-cuda` publishes an **opt-in** artifact and nothing activates it by accident. There is no +`META-INF/services` entry, so no `ServiceLoader` discovers it: a consumer has to call +`CudaGgufBatchedMatrixKernel.open()` and inject the kernel, and `-Dmodels.cuda.disabled=true` is a +kill switch on top of that. One `sm_80` PTX module serves every device of compute capability 8.0 or +above, carried in the jar under `META-INF/models/cuda/` with a SHA-256 the Java loader recomputes. + +**Its numeric state must be stated in the release notes, not assumed from its presence.** G4 (decode +speed) passes at 6.184x against a 3.00x gate, measured in +`benchmark-results/2026-10-07-g4-dualpath`. G1 (exact token parity against the CPU path) is a +separate gate with no tolerance, and a release note must say which way it last ran and on what +hardware. Shipping the jar does not mean the device path is numerically verified; it means a caller +can opt into it and read the gates. + ## Cut a release 1. Set a non-snapshot version in `gradle.properties`. @@ -23,10 +36,12 @@ published. API numeric kernels. The release workflow builds and tests the Models-owned Rust kernels on every supported native platform and compiles the Apple Foundation Models bridge on macOS before staging the signed Maven artifacts. -`backend-tornado` is an optional JVM artifact. The hosted release workflow verifies its Java -fallback and publication shape. Before release, run the public loader and exact CPU/GPU output -parity gate on each qualified NVIDIA hardware profile and retain the measurements under -`models-accelerator-bench/results/`. +Models ships **one** GPU implementation, `backend-cuda`. The TornadoVM arm was removed in 0.3.53: +measured on the same RTX 4090 and the same model, `backend-cuda` reached exact token parity and +31.26 tok/s decode where TornadoVM managed 6.49 tok/s over 36 failed `cuLaunchKernel` calls, and it +required a separately installed runtime that the Maven artifact could not carry. The release +precondition that pointed at `models-accelerator-bench/results/` went with it; `backend-cuda`'s +gates are `cuda-kernel-gate --mode capability|parity|decode` and need no external runtime. The workflow uses the same Maven Central and GPG secrets as `mfcqi-java`: `MAVENCENTRAL_USERNAME`, `MAVENCENTRAL_PASSWORD`, `GPG_PUBLIC_KEY`, diff --git a/backend-cuda/build.gradle.kts b/backend-cuda/build.gradle.kts index d2daddfd8..6318ce566 100644 --- a/backend-cuda/build.gradle.kts +++ b/backend-cuda/build.gradle.kts @@ -241,3 +241,32 @@ tasks.withType().configureEach { tasks.named("check") { dependsOn(cargoTestHost, cargoClippy, verifyPtxArtifact) } + +// Coverage: the 0.80 bar published modules carry applies to everything here that a host can +// execute, and two classes are exempted because they structurally cannot be. +// +// CudaDriver is the FFM binding to libcuda.so.1 -- every method is a downcall, so without a driver +// there is nothing to cover. CudaGgufBatchedMatrixKernel's bulk is the dispatch path behind those +// downcalls. Together they are 2,296 of the module's 2,491 missed instructions; the rest of the +// module measures 0.86 covered without them, and the classes a host *can* reach are already well +// past the bar -- CudaRoutingCounters at 391 of 403 instructions, Q8KActivations at 177 of 193. +// +// Their verification is hardware, not more off-device tests, and it is retained as evidence rather +// than asserted: G1 token parity and G4 decode speed in benchmark-results/2026-10-07-g1-parity and +// benchmark-results/2026-10-07-g4-dualpath, plus CudaQ6KDeviceParityTest, which launches the Q6_K +// kernel against the CPU control at six widths and skips off-device. Lowering the global bar to let +// this module in would have hidden a real gap in every other published module instead. +tasks.named("jacocoTestCoverageVerification") { + classDirectories.setFrom( + files( + classDirectories.files.map { directory -> + fileTree(directory) { + exclude( + "**/CudaDriver*.class", + "**/CudaGgufBatchedMatrixKernel*.class", + ) + } + }, + ), + ) +} diff --git a/backend-native/src/main/rust/model-kernels/Cargo.lock b/backend-native/src/main/rust/model-kernels/Cargo.lock index caf419773..117405e44 100644 --- a/backend-native/src/main/rust/model-kernels/Cargo.lock +++ b/backend-native/src/main/rust/model-kernels/Cargo.lock @@ -4,4 +4,4 @@ version = 4 [[package]] name = "jmodels-kernels" -version = "0.3.52" +version = "0.3.53" diff --git a/backend-native/src/main/rust/model-kernels/Cargo.toml b/backend-native/src/main/rust/model-kernels/Cargo.toml index 646fc0039..94bb99d18 100644 --- a/backend-native/src/main/rust/model-kernels/Cargo.toml +++ b/backend-native/src/main/rust/model-kernels/Cargo.toml @@ -1,6 +1,6 @@ [package] name = "jmodels-kernels" -version = "0.3.52" +version = "0.3.53" edition = "2024" license = "Apache-2.0" publish = false diff --git a/backend-tornado/.github/badges/mfcqi.json b/backend-tornado/.github/badges/mfcqi.json deleted file mode 100644 index 4b6b9094a..000000000 --- a/backend-tornado/.github/badges/mfcqi.json +++ /dev/null @@ -1,7 +0,0 @@ -{ - "schemaVersion": 1, - "label": "MFCQI", - "message": "0.83 (excellent)", - "color": "brightgreen", - "cacheSeconds": 3600 -} \ No newline at end of file diff --git a/backend-tornado/README.md b/backend-tornado/README.md deleted file mode 100644 index 180858449..000000000 --- a/backend-tornado/README.md +++ /dev/null @@ -1,35 +0,0 @@ -# Java GPU acceleration - -`backend-tornado` is an optional in-process device backend for Models. Its kernels are Java source; -TornadoVM compiles eligible Q4_0, Q4_K and Q6_K projections for a GPU at runtime. It does not invoke -an external model server or embed another inference engine. - -The Maven artifact does not bundle or transitively install TornadoVM's device runtime. Applications -provide a compatible TornadoVM distribution and launch configuration explicitly; without it, the -default loader uses the Java Vector API. - -Adding this module also registers its accelerator through Java `ServiceLoader`. -`PureJavaBackend.loadAutomatic(...)`—and ModelJars' Java backend path—select it only when the exact -artifact and device pass the capacity gate. - -The qualified production scope is deliberately narrow: - -- NVIDIA GPUs reached through TornadoVM's PTX backend; -- Q4_0 (Q8_0 activations) and Q4_K / Q6_K (Q8_K activations) GGUF projection work for both prefill - and single-token decode, with the two activation families never mixed inside one grouped dispatch; -- a single weight tensor under 2 GiB, the limit of TornadoVM's 32-bit array index; -- attention and unsupported tensor formats remain on the Java Vector API; and -- eager readiness compiles reusable plans before the first visible request. - -The capacity selector passed exact output-parity and full-model gates on NVIDIA A16 and A40 profiles -with Q4_0. The K-quant kernels are covered off-device against the vectors-core CPU kernels — exactly -on exactly-representable super-blocks, and within two float roundings per super-block otherwise — and -have not yet been run on a GPU. AMD, Intel, and Metal devices remain on the CPU fallback until they -pass equivalent real hardware gates. - -Run the hardware gate with the `accelerator-profile` command of `models-bench`, launched through the -TornadoVM launcher (see the models-bench README). It reports the selected device, readiness time, -prefill and decode throughput, and how many projections reached the device per GGUF weight format. - -See the published [Java GPU acceleration guide](https://integrallis.github.io/models/docs/models/current/gpu-acceleration.html) -for dependencies, launcher requirements, status reporting, and measured evidence. diff --git a/backend-tornado/build.gradle.kts b/backend-tornado/build.gradle.kts deleted file mode 100644 index aebb5933b..000000000 --- a/backend-tornado/build.gradle.kts +++ /dev/null @@ -1,55 +0,0 @@ -import org.gradle.testing.jacoco.tasks.JacocoCoverageVerification -import org.gradle.testing.jacoco.tasks.JacocoReport - -description = "Optional Java-authored GPU acceleration for the Models pure-Java backend" - -dependencies { - api(project(":backend-java")) - compileOnly("io.github.beehive-lab:tornado-api:5.2.0-jdk25") - testImplementation("io.github.beehive-lab:tornado-api:5.2.0-jdk25") - testRuntimeOnly("io.github.beehive-lab:tornado-runtime:5.2.0-jdk25") - testImplementation("com.integrallis:vectors-core:${providers.gradleProperty("vectorsVersion").get()}") -} - -val configuredQwenModel = providers.systemProperty("models.fixtures.qwen306BQ40") -val acceleratorRequired = providers.systemProperty("models.accelerator.required") -val acceleratorExpected = providers.systemProperty("models.accelerator.expected") - -tasks.withType().configureEach { - configuredQwenModel.orNull?.let { - systemProperty("models.fixtures.qwen306BQ40", it) - } - acceleratorRequired.orNull?.let { - systemProperty("models.accelerator.required", it) - } - acceleratorExpected.orNull?.let { - systemProperty("models.accelerator.expected", it) - } -} - -// Driver discovery, model loading, and Tornado execution plans are covered by the opt-in model -// integration test and the release hardware gates. Keep the ordinary unit-coverage denominator on -// the device-independent selection, validation, and Java kernel logic. -val hardwareIntegrationClasses = - listOf( - "**/TornadoBackend.class", - "**/TornadoBackendRuntime.class", - "**/TornadoRuntimeDevices.class", - "**/TornadoGgufBatchedMatrixKernel*.class" - ) - -tasks.withType().configureEach { - classDirectories.setFrom( - files(classDirectories.files.map { directory -> - fileTree(directory) { exclude(hardwareIntegrationClasses) } - }) - ) -} - -tasks.withType().configureEach { - classDirectories.setFrom( - files(classDirectories.files.map { directory -> - fileTree(directory) { exclude(hardwareIntegrationClasses) } - }) - ) -} diff --git a/backend-tornado/gradle.lockfile b/backend-tornado/gradle.lockfile deleted file mode 100644 index f8ef23f52..000000000 --- a/backend-tornado/gradle.lockfile +++ /dev/null @@ -1,75 +0,0 @@ -# This is a Gradle generated file for dependency locking. -# Manual edits can break the build and are not advised. -# This file is expected to be part of source control. -biz.aQute.bnd:biz.aQute.bnd.annotation:7.1.0=compileClasspath,testCompileClasspath -com.fasterxml.jackson.core:jackson-core:2.21.7=runtimeClasspath,testRuntimeClasspath -com.fasterxml.jackson:jackson-bom:2.21.7=runtimeClasspath,testRuntimeClasspath -com.github.spotbugs:spotbugs-annotations:4.9.8=spotbugs -com.github.spotbugs:spotbugs:4.9.8=spotbugs -com.github.stephenc.jcip:jcip-annotations:1.0-1=spotbugs -com.google.code.findbugs:jsr305:3.0.2=spotbugs -com.google.code.gson:gson:2.13.2=spotbugs -com.google.errorprone:error_prone_annotations:2.38.0=compileClasspath,testCompileClasspath -com.google.errorprone:error_prone_annotations:2.41.0=spotbugs -com.integrallis:vectors-core:0.1.28=runtimeClasspath,testCompileClasspath,testRuntimeClasspath -commons-io:commons-io:2.20.0=spotbugs -io.github.beehive-lab:tornado-api:5.2.0-jdk25=compileClasspath,testCompileClasspath,testRuntimeClasspath -io.github.beehive-lab:tornado-runtime:5.2.0-jdk25=testRuntimeClasspath -jaxen:jaxen:2.0.0=spotbugs -net.bytebuddy:byte-buddy-agent:1.17.7=testCompileClasspath,testRuntimeClasspath -net.bytebuddy:byte-buddy:1.17.7=testCompileClasspath,testRuntimeClasspath -net.sf.jopt-simple:jopt-simple:4.6=compileClasspath,testCompileClasspath,testRuntimeClasspath -net.sf.saxon:Saxon-HE:12.9=spotbugs -org.apache.bcel:bcel:6.11.0=spotbugs -org.apache.commons:commons-lang3:3.19.0=spotbugs -org.apache.commons:commons-math3:3.2=compileClasspath,testCompileClasspath,testRuntimeClasspath -org.apache.commons:commons-text:1.14.0=spotbugs -org.apache.logging.log4j:log4j-api:2.25.2=spotbugs -org.apache.logging.log4j:log4j-api:2.25.4=compileClasspath,testCompileClasspath,testRuntimeClasspath -org.apache.logging.log4j:log4j-core:2.25.2=spotbugs -org.apache.logging.log4j:log4j-core:2.25.4=compileClasspath,testCompileClasspath,testRuntimeClasspath -org.apiguardian:apiguardian-api:1.1.2=testCompileClasspath -org.assertj:assertj-core:3.27.2=testCompileClasspath,testRuntimeClasspath -org.dom4j:dom4j:2.2.0=spotbugs -org.graalvm.polyglot:polyglot:25.0.2=compileClasspath,testCompileClasspath,testRuntimeClasspath -org.graalvm.sdk:collections:25.0.2=compileClasspath,testCompileClasspath,testRuntimeClasspath -org.graalvm.sdk:nativeimage:25.0.2=compileClasspath,testCompileClasspath,testRuntimeClasspath -org.graalvm.sdk:word:25.0.2=compileClasspath,testCompileClasspath,testRuntimeClasspath -org.jacoco:org.jacoco.agent:0.8.14=jacocoAgent,jacocoAnt -org.jacoco:org.jacoco.ant:0.8.14=jacocoAnt -org.jacoco:org.jacoco.core:0.8.14=jacocoAnt -org.jacoco:org.jacoco.report:0.8.14=jacocoAnt -org.jspecify:jspecify:1.0.0=compileClasspath,testCompileClasspath -org.junit.jupiter:junit-jupiter-api:5.11.4=testCompileClasspath -org.junit.jupiter:junit-jupiter-api:5.13.4=testRuntimeClasspath -org.junit.jupiter:junit-jupiter-engine:5.13.4=testRuntimeClasspath -org.junit.jupiter:junit-jupiter-params:5.11.4=testCompileClasspath -org.junit.jupiter:junit-jupiter-params:5.13.4=testRuntimeClasspath -org.junit.jupiter:junit-jupiter:5.11.4=testCompileClasspath -org.junit.jupiter:junit-jupiter:5.13.4=testRuntimeClasspath -org.junit.platform:junit-platform-commons:1.11.4=testCompileClasspath -org.junit.platform:junit-platform-commons:1.13.4=testRuntimeClasspath -org.junit.platform:junit-platform-engine:1.13.4=testRuntimeClasspath -org.junit.platform:junit-platform-launcher:1.13.4=testRuntimeClasspath -org.junit:junit-bom:5.11.4=testCompileClasspath -org.junit:junit-bom:5.13.4=testRuntimeClasspath -org.junit:junit-bom:5.14.0=spotbugs -org.mockito:mockito-core:5.23.0=testCompileClasspath,testRuntimeClasspath -org.mockito:mockito-junit-jupiter:5.23.0=testCompileClasspath,testRuntimeClasspath -org.objenesis:objenesis:3.3=testRuntimeClasspath -org.openjdk.jmh:jmh-core:1.29=compileClasspath,testCompileClasspath,testRuntimeClasspath -org.opentest4j:opentest4j:1.3.0=testCompileClasspath,testRuntimeClasspath -org.osgi:org.osgi.annotation.bundle:2.0.0=compileClasspath,testCompileClasspath -org.osgi:org.osgi.annotation.versioning:1.1.2=compileClasspath,testCompileClasspath -org.osgi:org.osgi.resource:1.0.0=compileClasspath,testCompileClasspath -org.osgi:org.osgi.service.serviceloader:1.0.0=compileClasspath,testCompileClasspath -org.ow2.asm:asm-analysis:9.9=spotbugs -org.ow2.asm:asm-commons:9.9=jacocoAnt,spotbugs -org.ow2.asm:asm-tree:9.9=jacocoAnt,spotbugs -org.ow2.asm:asm-util:9.9=spotbugs -org.ow2.asm:asm:9.9=jacocoAnt,spotbugs -org.slf4j:slf4j-api:2.0.17=runtimeClasspath,spotbugs,spotbugsSlf4j,testRuntimeClasspath -org.slf4j:slf4j-simple:2.0.17=spotbugsSlf4j -org.snmp4j:snmp4j:2.8.6=testRuntimeClasspath -org.xmlresolver:xmlresolver:5.3.3=spotbugs -empty=annotationProcessor,cyclonedxBom,spotbugsPlugins,testAnnotationProcessor diff --git a/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/AcceleratorEligibility.java b/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/AcceleratorEligibility.java deleted file mode 100644 index da1788871..000000000 --- a/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/AcceleratorEligibility.java +++ /dev/null @@ -1,320 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.backend.tornado; - -import java.time.Duration; -import java.util.Comparator; -import java.util.List; -import java.util.Locale; -import java.util.Objects; - -/** - * Device-vendor and capacity policy applied before constructing accelerator execution plans. - * - *

The gate adds up what a model would actually place on the device — retained weights under the - * plan shape the kernel builds, per-plan scratch, and any device-resident KV cache — and compares - * it against the device's global memory less a safety margin. Where a term is unknown it is refused - * rather than assumed to be zero: a model whose weights exceed {@link - * #COARSE_PROFILE_WEIGHT_LIMIT_BYTES} is not admitted on a file-size-only budget, because at that - * size the unknown terms are larger than the margin. - */ -public final class AcceleratorEligibility { - - /** - * Flat allowance for the TornadoVM device context, compiled code, and driver bookkeeping. - * - *

Unmeasured. It is a constant carried over from the published selector, kept so the admitted - * budget for the qualified A16-2Q and A40-4Q profiles is unchanged. Calibrating it needs a GPU - * host: {@code TornadoExecutionPlan.getCurrentDeviceMemoryUsage()} reports real usage after - * readiness, and until that is run this term must not be described as measured. - */ - static final long BASE_PLAN_OVERHEAD_BYTES = 256L * 1024L * 1024L; - - /** Maximum single allocation assumed for a coarse, file-size-only request. */ - static final long COARSE_MAX_SINGLE_ALLOCATION_BYTES = 512L * 1024L * 1024L; - - /** - * Hard cap on one uploaded tensor. - * - *

{@code ByteArray} stores its element count in an {@code int} and {@code - * ByteArray.fromSegment} computes it as {@code (int) segment.byteSize()}, so a tensor at or above - * 2 GiB is truncated on construction and the following {@code MemorySegment.copy} of the full - * tensor fails. Read from {@code tornado-api 5.2.0-jdk25}. - */ - static final long MAX_TORNADO_ARRAY_BYTES = Integer.MAX_VALUE; - - /** - * Weight ceiling for admitting a model on a coarse, file-size-only request. - * - *

Below this a model's KV cache, per-plan scratch, and largest tensor all fit comfortably - * inside {@link #BASE_PLAN_OVERHEAD_BYTES}; above it they do not, and admitting the model would - * mean gating on a number known to be incomplete. - */ - static final long COARSE_PROFILE_WEIGHT_LIMIT_BYTES = 8L * 1024L * 1024L * 1024L; - - /** - * Eager-readiness ceiling above which a configuration is not deployable. - * - *

Fixed in advance by gate G5 of the large-model GPU campaign pre-registration. - */ - static final Duration READINESS_BUDGET = Duration.ofSeconds(120); - - /** Fraction of device global memory the gate refuses to plan into. */ - private static final long SAFETY_DIVISOR = 4L; - - private AcceleratorEligibility() {} - - /** - * Selects a device using the coarse, file-size-only budget available before a model is parsed. - * - * @param devices devices reported by the TornadoVM runtime - * @param modelSizeBytes size of the GGUF file on disk - * @param accelerateDecode whether separate single-token decode plans are retained - */ - public static Decision select( - List devices, long modelSizeBytes, boolean accelerateDecode) { - Objects.requireNonNull(devices, "devices"); - if (modelSizeBytes <= 0) { - return Decision.ineligible("model size must be positive", 0, null); - } - return select( - devices, DeviceMemoryRequest.ofModelFile("model", modelSizeBytes, accelerateDecode)); - } - - /** Selects a device for a fully specified device-memory request. */ - public static Decision select(List devices, DeviceMemoryRequest request) { - Objects.requireNonNull(devices, "devices"); - Objects.requireNonNull(request, "request"); - - DeviceBudget shipped = budget(request, PlanShapeStrategy.PER_SHAPE_WHOLE_MODEL); - - List qualified = - devices.stream() - .filter(device -> "PTX".equals(device.backend()) || "CUDA".equals(device.backend())) - .filter(device -> "GPU".equals(device.type())) - .sorted(Comparator.comparingLong(DeviceCapabilities::globalMemoryBytes).reversed()) - .toList(); - if (qualified.isEmpty()) { - return Decision.ineligible( - "no qualified NVIDIA GPU backend was discovered; needed a PTX or CUDA backend on a GPU" - + " device, found " - + describe(devices), - shipped.totalBytes(), - shipped); - } - - if (request.largestAllocationBytes() >= MAX_TORNADO_ARRAY_BYTES) { - return Decision.ineligible( - "largest device allocation is " - + DeviceBudget.gib(request.largestAllocationBytes()) - + " but a TornadoVM ByteArray holds at most " - + DeviceBudget.gib(MAX_TORNADO_ARRAY_BYTES) - + " (int element count); the tensor must be split before it can be uploaded", - shipped.totalBytes(), - shipped); - } - - if (!request.detailed() && request.weightBytes() > COARSE_PROFILE_WEIGHT_LIMIT_BYTES) { - return Decision.ineligible( - "model '" - + request.modelLabel() - + "' has " - + DeviceBudget.gib(request.weightBytes()) - + " of weights, above the " - + DeviceBudget.gib(COARSE_PROFILE_WEIGHT_LIMIT_BYTES) - + " limit for a file-size-only device budget; KV residency, retained-plan count," - + " per-plan scratch and largest-tensor size are all unknown in this request." - + " Build a DeviceMemoryRequest.detailed(...) from the model's tensor inventory" - + " before admitting a model this size", - shipped.totalBytes(), - shipped); - } - - DeviceCapabilities best = qualified.get(0); - for (DeviceCapabilities device : qualified) { - long safeCapacity = safeCapacity(device); - if (shipped.totalBytes() <= safeCapacity && allocationFits(request, device)) { - Duration readiness = shipped.estimatedPlanCompileTime(); - if (readiness.compareTo(READINESS_BUDGET) > 0) { - return Decision.ineligible( - "device " - + device.name() - + " fits the plan but the " - + request.retainedPlanCount() - + " retained plans would need about " - + seconds(readiness) - + " of eager readiness at the measured " - + String.format(Locale.ROOT, "%.1f", DeviceBudget.PLANS_COMPILED_PER_SECOND) - + " plans/s (A40-4Q, 223 plans in 13.554 s), above the " - + seconds(READINESS_BUDGET) - + " deployable ceiling; reduce the retained plan count before admitting this" - + " model", - shipped.totalBytes(), - shipped); - } - return new Decision( - true, - device, - "eligible", - shipped.totalBytes(), - shipped, - PlanShapeStrategy.PER_SHAPE_WHOLE_MODEL); - } - } - - for (DeviceCapabilities device : qualified) { - if (!allocationFits(request, device)) { - continue; - } - long safeCapacity = safeCapacity(device); - for (PlanShapeStrategy strategy : PlanShapeStrategy.values()) { - if (strategy.implemented() || !strategy.capacityComputable()) { - continue; - } - DeviceBudget alternative = budget(request, strategy); - if (alternative.totalBytes() <= safeCapacity) { - return Decision.ineligible( - "device " - + device.name() - + " has " - + DeviceBudget.gib(safeCapacity) - + " usable of " - + DeviceBudget.gib(device.globalMemoryBytes()) - + " but the shipped plan shape needs " - + DeviceBudget.gib(shipped.totalBytes()) - + " (" - + shipped.describe() - + "); the model would fit in " - + DeviceBudget.gib(alternative.totalBytes()) - + " under " - + strategy - + ", which is not implemented: " - + strategy.limitation(), - shipped.totalBytes(), - shipped); - } - } - } - - DeviceCapabilities target = best; - long safeCapacity = safeCapacity(target); - if (!allocationFits(request, target)) { - return Decision.ineligible( - "device " - + target.name() - + " allows a single allocation of at most " - + DeviceBudget.gib(target.maxAllocationBytes()) - + " but the plan needs one of " - + DeviceBudget.gib(largestAllocation(request)), - shipped.totalBytes(), - shipped); - } - return Decision.ineligible( - "device " - + target.name() - + " has " - + DeviceBudget.gib(safeCapacity) - + " usable of " - + DeviceBudget.gib(target.globalMemoryBytes()) - + " but the plan needs " - + DeviceBudget.gib(shipped.totalBytes()) - + " (" - + shipped.describe() - + "), short by " - + DeviceBudget.gib(shipped.totalBytes() - safeCapacity) - + "; the same plans also hold " - + DeviceBudget.gib(shipped.hostWeightCopyBytes()) - + " of host weight copies", - shipped.totalBytes(), - shipped); - } - - /** Computes the device budget a request costs under one plan shape. */ - public static DeviceBudget budget(DeviceMemoryRequest request, PlanShapeStrategy strategy) { - return DeviceBudget.of(request, strategy, BASE_PLAN_OVERHEAD_BYTES); - } - - private static long safeCapacity(DeviceCapabilities device) { - return device.globalMemoryBytes() - device.globalMemoryBytes() / SAFETY_DIVISOR; - } - - private static long largestAllocation(DeviceMemoryRequest request) { - return request.detailed() - ? request.largestAllocationBytes() - : Math.min(request.weightBytes(), COARSE_MAX_SINGLE_ALLOCATION_BYTES); - } - - private static boolean allocationFits(DeviceMemoryRequest request, DeviceCapabilities device) { - return largestAllocation(request) <= device.maxAllocationBytes(); - } - - private static String describe(List devices) { - if (devices.isEmpty()) { - return "no devices"; - } - StringBuilder text = new StringBuilder(64); - for (DeviceCapabilities device : devices) { - if (text.length() > 0) { - text.append(", "); - } - text.append(device.name()) - .append(" [") - .append(device.backend()) - .append('/') - .append(device.type()) - .append(']'); - } - return text.toString(); - } - - private static String seconds(Duration duration) { - return String.format(Locale.ROOT, "%.0f s", duration.toMillis() / 1000.0); - } - - /** One device as reported by the TornadoVM runtime. */ - public record DeviceCapabilities( - String name, String backend, String type, long globalMemoryBytes, long maxAllocationBytes) { - public DeviceCapabilities { - Objects.requireNonNull(name, "name"); - Objects.requireNonNull(backend, "backend"); - Objects.requireNonNull(type, "type"); - } - } - - /** - * The outcome of device selection. - * - * @param eligible whether a device was admitted - * @param device the admitted device, or {@code null} - * @param reason {@code "eligible"}, or a message naming what was needed and what was available - * @param requiredBytes device bytes the shipped plan shape would need - * @param budget the itemised budget behind {@code requiredBytes}, or {@code null} when the - * request was rejected before one could be formed - * @param strategy the plan shape the admitted device was admitted under, or {@code null} - */ - public record Decision( - boolean eligible, - DeviceCapabilities device, - String reason, - long requiredBytes, - DeviceBudget budget, - PlanShapeStrategy strategy) { - - private static Decision ineligible(String reason, long requiredBytes, DeviceBudget budget) { - return new Decision(false, null, reason, requiredBytes, budget, null); - } - } -} diff --git a/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/DeviceBudget.java b/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/DeviceBudget.java deleted file mode 100644 index ff87a937b..000000000 --- a/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/DeviceBudget.java +++ /dev/null @@ -1,101 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.backend.tornado; - -import java.time.Duration; -import java.util.Locale; -import java.util.Objects; - -/** - * The device and host bytes one {@link DeviceMemoryRequest} costs under one {@link - * PlanShapeStrategy}, itemised so an ineligibility message can name every term. - * - * @param strategy the plan shape this budget was computed for - * @param deviceWeightBytes model weights retained on the device - * @param planScratchBytes device activation, scale, and output buffers across all retained plans - * @param kvCacheBytes KV cache bytes retained on the device - * @param baseOverheadBytes flat runtime allowance for the TornadoVM context and compiled code - * @param totalBytes sum of the device terms - * @param hostWeightCopyBytes off-heap host copies the plans hold, on top of the mapped GGUF - * @param estimatedPlanCompileTime eager-readiness estimate from the measured plan compile rate - */ -public record DeviceBudget( - PlanShapeStrategy strategy, - long deviceWeightBytes, - long planScratchBytes, - long kvCacheBytes, - long baseOverheadBytes, - long totalBytes, - long hostWeightCopyBytes, - Duration estimatedPlanCompileTime) { - - /** - * Plans compiled per second during eager readiness. - * - *

Measured, not assumed: the Vultr A40-4Q gate on 2026-08-29 compiled 223 retained plans in - * 13.554 s ({@code models-accelerator-bench/results/vultr-a40-4q-2026-08-29.md}). The A16-2Q gate - * on the same artifact took 14.278 s for the same plan set, so this rate is the faster of the two - * measured profiles and the readiness estimate it produces is optimistic. - */ - static final double PLANS_COMPILED_PER_SECOND = 223.0 / 13.554; - - public DeviceBudget { - Objects.requireNonNull(strategy, "strategy"); - Objects.requireNonNull(estimatedPlanCompileTime, "estimatedPlanCompileTime"); - } - - static DeviceBudget of(DeviceMemoryRequest request, PlanShapeStrategy strategy, long baseBytes) { - Objects.requireNonNull(request, "request"); - Objects.requireNonNull(strategy, "strategy"); - long weights = strategy.deviceWeightBytes(request); - long scratch = request.planScratchBytes(); - long kv = request.deviceKvCacheBytes(); - long total = Math.addExact(Math.addExact(weights, scratch), Math.addExact(kv, baseBytes)); - Duration readiness = - Duration.ofMillis( - Math.round(request.retainedPlanCount() / PLANS_COMPILED_PER_SECOND * 1000.0)); - return new DeviceBudget( - strategy, - weights, - scratch, - kv, - baseBytes, - total, - strategy.hostWeightCopyBytes(request), - readiness); - } - - /** A one-line itemisation suitable for an operator-facing message. */ - public String describe() { - StringBuilder text = new StringBuilder(128); - text.append("weights ") - .append(gib(deviceWeightBytes)) - .append(" (") - .append(strategy) - .append("), KV ") - .append(gib(kvCacheBytes)) - .append(", plan scratch ") - .append(gib(planScratchBytes)) - .append(", base ") - .append(gib(baseOverheadBytes)); - return text.toString(); - } - - /** Formats a byte count as GiB with two decimals. */ - public static String gib(long bytes) { - return String.format(Locale.ROOT, "%.2f GiB", bytes / (double) (1024L * 1024L * 1024L)); - } -} diff --git a/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/DeviceMemoryRequest.java b/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/DeviceMemoryRequest.java deleted file mode 100644 index 8cb69c789..000000000 --- a/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/DeviceMemoryRequest.java +++ /dev/null @@ -1,179 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.backend.tornado; - -import java.util.Objects; - -/** - * What one model would place on an accelerator, in the terms the capacity gate can add up. - * - *

A request is either coarse or detailed. A coarse request is derived from the - * GGUF file size alone: it knows the upper bound on uploaded weights and nothing else, so retained - * plan count, largest single tensor, and device KV residency are all reported as unknown rather - * than silently assumed to be zero. A detailed request is built from the tensor inventory and - * carries all of those terms. - * - * @param modelLabel human-readable model identity used in ineligibility messages - * @param weightBytes bytes of model weights that would be uploaded once, per retained shape - * @param retainedShapes distinct batch shapes whose plans are held at once (prefill, or prefill and - * decode) - * @param retainedPlanCount distinct {@code (weight, shape)} plans that would be retained; {@code 0} - * means unknown - * @param planScratchBytes device activation, scale, and output buffers summed over every retained - * plan; {@code 0} means unknown - * @param largestAllocationBytes largest single device allocation the plans would ask for, whether a - * weight tensor or a scratch buffer; {@code 0} means unknown - * @param deviceKvCacheBytes KV cache bytes that would live on the device at the planned context - * length; {@code 0} means the KV cache stays in host memory - * @param detailed whether the weight, plan, and tensor terms came from a real tensor inventory - */ -public record DeviceMemoryRequest( - String modelLabel, - long weightBytes, - int retainedShapes, - int retainedPlanCount, - long planScratchBytes, - long largestAllocationBytes, - long deviceKvCacheBytes, - boolean detailed) { - - public DeviceMemoryRequest { - modelLabel = requireLabel(modelLabel); - requirePositive("weightBytes", weightBytes); - if (retainedShapes < 1) { - throw new IllegalArgumentException("retainedShapes must be at least 1: " + retainedShapes); - } - requireNotNegative("retainedPlanCount", retainedPlanCount); - requireNotNegative("planScratchBytes", planScratchBytes); - requireNotNegative("largestAllocationBytes", largestAllocationBytes); - requireNotNegative("deviceKvCacheBytes", deviceKvCacheBytes); - } - - /** - * Builds the coarse request the loader can form before parsing a model: the whole GGUF file is - * treated as uploadable weights and nothing else is known. - * - *

This is what {@code TornadoBackend.open} has available, and it reproduces the budget the - * published A16-2Q and A40-4Q qualification runs were admitted under. - */ - public static DeviceMemoryRequest ofModelFile( - String modelLabel, long fileBytes, boolean accelerateDecode) { - return new DeviceMemoryRequest( - modelLabel, fileBytes, accelerateDecode ? 2 : 1, 0, 0L, 0L, 0L, false); - } - - /** Starts a detailed request built from a real tensor inventory. */ - public static Builder detailed(String modelLabel) { - return new Builder(modelLabel); - } - - /** Returns this request with a device-resident KV cache of the given size. */ - public DeviceMemoryRequest withDeviceKvCacheBytes(long bytes) { - return new DeviceMemoryRequest( - modelLabel, - weightBytes, - retainedShapes, - retainedPlanCount, - planScratchBytes, - largestAllocationBytes, - bytes, - detailed); - } - - /** Mutable assembly for a detailed request. */ - public static final class Builder { - private final String modelLabel; - private long weightBytes; - private int retainedShapes = 1; - private int retainedPlanCount; - private long planScratchBytes; - private long largestAllocationBytes; - private long deviceKvCacheBytes; - - private Builder(String modelLabel) { - this.modelLabel = requireLabel(modelLabel); - } - - /** Sets the bytes of weights uploaded once, per retained shape. */ - public Builder weightBytes(long bytes) { - this.weightBytes = bytes; - return this; - } - - /** Sets how many batch shapes are retained at once. */ - public Builder retainedShapes(int shapes) { - this.retainedShapes = shapes; - return this; - } - - /** Sets how many distinct {@code (weight, shape)} plans are retained. */ - public Builder retainedPlanCount(int plans) { - this.retainedPlanCount = plans; - return this; - } - - /** Sets the device activation, scale, and output buffers summed over every retained plan. */ - public Builder planScratchBytes(long bytes) { - this.planScratchBytes = bytes; - return this; - } - - /** Sets the largest single device allocation the plans would ask for. */ - public Builder largestAllocationBytes(long bytes) { - this.largestAllocationBytes = bytes; - return this; - } - - /** Sets the KV cache bytes that would live on the device at the planned context length. */ - public Builder deviceKvCacheBytes(long bytes) { - this.deviceKvCacheBytes = bytes; - return this; - } - - /** Builds the detailed request. */ - public DeviceMemoryRequest build() { - return new DeviceMemoryRequest( - modelLabel, - weightBytes, - retainedShapes, - retainedPlanCount, - planScratchBytes, - largestAllocationBytes, - deviceKvCacheBytes, - true); - } - } - - private static String requireLabel(String value) { - Objects.requireNonNull(value, "modelLabel"); - if (value.isBlank()) { - throw new IllegalArgumentException("modelLabel must not be blank"); - } - return value.strip(); - } - - private static void requirePositive(String name, long value) { - if (value <= 0) { - throw new IllegalArgumentException(name + " must be positive: " + value); - } - } - - private static void requireNotNegative(String name, long value) { - if (value < 0) { - throw new IllegalArgumentException(name + " must not be negative: " + value); - } - } -} diff --git a/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/KQuantProjectionKernel.java b/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/KQuantProjectionKernel.java deleted file mode 100644 index 29cf10ab7..000000000 --- a/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/KQuantProjectionKernel.java +++ /dev/null @@ -1,568 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.backend.tornado; - -import uk.ac.manchester.tornado.api.annotations.Parallel; -import uk.ac.manchester.tornado.api.types.arrays.ByteArray; -import uk.ac.manchester.tornado.api.types.arrays.FloatArray; -import uk.ac.manchester.tornado.api.types.arrays.IntArray; - -/** - * Java-authored TornadoVM kernels for production-compatible K-quant by Q8_K projections. - * - *

Q4_K_M catalog models mix two super-block formats in one tensor set: Q4_K for most projections - * and Q6_K for the value and a share of the down projections. Both consume the same Q8_K activation - * preparation, so one host-side quantization feeds every kernel here. - * - *

GGUF super-block layouts implemented here (little-endian, 256 values per super-block): - * - *

    - *
  • Q4_K, 144 bytes: {@code d} (fp16) at 0, {@code dmin} (fp16) at 2, twelve bytes of - * six-bit packed scales and mins at 4, and 128 bytes of 4-bit quants at 16. The 256 values - * are eight groups of 32; group {@code g} reads the 32 bytes at {@code 16 + (g >>> 1) * 32} - * and takes the low nibble when {@code g} is even and the high nibble when it is odd. Group - * values are unsigned 0..15 and the offset is carried by the per-group six-bit minimum. - *
  • Q6_K, 210 bytes: 128 bytes of low nibbles, 64 bytes of high two-bit pairs, 16 signed - * int8 group scales, then {@code d} (fp16) at 208. Quants are {@code (ql | qh << 4) - 32}. - *
- * - *

Arithmetic contract. Every per-super-block reduction is accumulated in {@code int}, so - * the quantized dot products and the Q4_K minimum corrections are exact and independent of work - * scheduling. Only the per-super-block scale application is floating point, and it is applied in - * ascending super-block order with the same two-step form the vectors-core CPU kernels use ({@code - * sum + d * quantizedSum}, then {@code sum - dMin * minimumSum}). The CPU kernels fuse those two - * steps with {@code Math.fma}; this kernel uses a plain multiply and add, matching the operation - * set the existing Q4_0 kernel is known to compile to PTX with. The two agree bit-for-bit whenever - * the products are exactly representable and differ by at most one fused-multiply-add rounding per - * super-block otherwise. - */ -public final class KQuantProjectionKernel { - /** Values per K-quant super-block. */ - public static final int SUPER_BLOCK_VALUES = 256; - - /** Values per Q8_K partial-sum block; Q4_K's minimum correction reads two of these per group. */ - public static final int SUM_BLOCK_VALUES = 16; - - /** Q8_K partial-sum entries per super-block. */ - public static final int SUMS_PER_SUPER_BLOCK = SUPER_BLOCK_VALUES / SUM_BLOCK_VALUES; - - /** Bytes per Q4_K super-block. */ - public static final int Q4_K_BLOCK_BYTES = 144; - - /** Bytes per Q6_K super-block. */ - public static final int Q6_K_BLOCK_BYTES = 210; - - private static final int Q4_K_SCALES_OFFSET = 4; - private static final int Q4_K_QUANTS_OFFSET = 16; - private static final int Q4_K_GROUPS = 8; - private static final int Q4_K_GROUP_VALUES = 32; - - private static final int Q6_K_QL_BYTES = 128; - private static final int Q6_K_QH_BYTES = 64; - private static final int Q6_K_SCALE_BYTES = 16; - private static final int Q6_K_SCALE_OFFSET = Q6_K_QL_BYTES + Q6_K_QH_BYTES; - private static final int Q6_K_DELTA_OFFSET = Q6_K_SCALE_OFFSET + Q6_K_SCALE_BYTES; - - private KQuantProjectionKernel() {} - - /** Runs one work item per batch/output-row pair over host-prepared Q8_K activations. */ - public static void multiplyQ4K( - ByteArray weights, - ByteArray activations, - FloatArray activationScales, - IntArray activationSums, - FloatArray output, - int batchSize, - int rows, - int cols) { - int blocks = cols / SUPER_BLOCK_VALUES; - int sumsPerBatch = cols / SUM_BLOCK_VALUES; - for (@Parallel int outputIndex = 0; outputIndex < batchSize * rows; outputIndex++) { - int batch = outputIndex / rows; - int row = outputIndex - batch * rows; - output.set( - outputIndex, - q4kRowDot( - weights, - row, - activations, - batch * cols, - activationScales, - batch * blocks, - activationSums, - batch * sumsPerBatch, - blocks)); - } - } - - /** Computes two Q4_K projections in one dispatch over one prepared activation. */ - public static void multiplyQ4KDual( - ByteArray firstWeights, - int firstRows, - ByteArray secondWeights, - int secondRows, - ByteArray activations, - FloatArray activationScales, - IntArray activationSums, - FloatArray firstOutput, - FloatArray secondOutput, - int batchSize, - int cols) { - int blocks = cols / SUPER_BLOCK_VALUES; - int sumsPerBatch = cols / SUM_BLOCK_VALUES; - int combinedRows = firstRows + secondRows; - for (@Parallel int outputIndex = 0; outputIndex < batchSize * combinedRows; outputIndex++) { - int batch = outputIndex / combinedRows; - int combinedRow = outputIndex - batch * combinedRows; - if (combinedRow < firstRows) { - firstOutput.set( - batch * firstRows + combinedRow, - q4kRowDot( - firstWeights, - combinedRow, - activations, - batch * cols, - activationScales, - batch * blocks, - activationSums, - batch * sumsPerBatch, - blocks)); - } else { - int row = combinedRow - firstRows; - secondOutput.set( - batch * secondRows + row, - q4kRowDot( - secondWeights, - row, - activations, - batch * cols, - activationScales, - batch * blocks, - activationSums, - batch * sumsPerBatch, - blocks)); - } - } - } - - /** Computes three Q4_K projections in one dispatch over one prepared activation. */ - public static void multiplyQ4KTriple( - ByteArray firstWeights, - int firstRows, - ByteArray secondWeights, - int secondRows, - ByteArray thirdWeights, - int thirdRows, - ByteArray activations, - FloatArray activationScales, - IntArray activationSums, - FloatArray firstOutput, - FloatArray secondOutput, - FloatArray thirdOutput, - int batchSize, - int cols) { - int blocks = cols / SUPER_BLOCK_VALUES; - int sumsPerBatch = cols / SUM_BLOCK_VALUES; - int firstAndSecondRows = firstRows + secondRows; - int combinedRows = firstAndSecondRows + thirdRows; - for (@Parallel int outputIndex = 0; outputIndex < batchSize * combinedRows; outputIndex++) { - int batch = outputIndex / combinedRows; - int combinedRow = outputIndex - batch * combinedRows; - if (combinedRow < firstRows) { - firstOutput.set( - batch * firstRows + combinedRow, - q4kRowDot( - firstWeights, - combinedRow, - activations, - batch * cols, - activationScales, - batch * blocks, - activationSums, - batch * sumsPerBatch, - blocks)); - } else if (combinedRow < firstAndSecondRows) { - int row = combinedRow - firstRows; - secondOutput.set( - batch * secondRows + row, - q4kRowDot( - secondWeights, - row, - activations, - batch * cols, - activationScales, - batch * blocks, - activationSums, - batch * sumsPerBatch, - blocks)); - } else { - int row = combinedRow - firstAndSecondRows; - thirdOutput.set( - batch * thirdRows + row, - q4kRowDot( - thirdWeights, - row, - activations, - batch * cols, - activationScales, - batch * blocks, - activationSums, - batch * sumsPerBatch, - blocks)); - } - } - } - - /** Runs one work item per batch/output-row pair over host-prepared Q8_K activations. */ - public static void multiplyQ6K( - ByteArray weights, - ByteArray activations, - FloatArray activationScales, - FloatArray output, - int batchSize, - int rows, - int cols) { - int blocks = cols / SUPER_BLOCK_VALUES; - for (@Parallel int outputIndex = 0; outputIndex < batchSize * rows; outputIndex++) { - int batch = outputIndex / rows; - int row = outputIndex - batch * rows; - output.set( - outputIndex, - q6kRowDot( - weights, row, activations, batch * cols, activationScales, batch * blocks, blocks)); - } - } - - /** - * Computes the Q4_K/Q4_K/Q6_K attention projection group in one dispatch. - * - *

This is the shape a Q4_K_M model presents for query, key, and value: llama.cpp promotes the - * value projection to Q6_K while the query and key projections stay Q4_K. - */ - public static void multiplyMixedTriple( - ByteArray firstWeights, - int firstRows, - ByteArray secondWeights, - int secondRows, - ByteArray thirdWeights, - int thirdRows, - ByteArray activations, - FloatArray activationScales, - IntArray activationSums, - FloatArray firstOutput, - FloatArray secondOutput, - FloatArray thirdOutput, - int batchSize, - int cols) { - int blocks = cols / SUPER_BLOCK_VALUES; - int sumsPerBatch = cols / SUM_BLOCK_VALUES; - int firstAndSecondRows = firstRows + secondRows; - int combinedRows = firstAndSecondRows + thirdRows; - for (@Parallel int outputIndex = 0; outputIndex < batchSize * combinedRows; outputIndex++) { - int batch = outputIndex / combinedRows; - int combinedRow = outputIndex - batch * combinedRows; - if (combinedRow < firstRows) { - firstOutput.set( - batch * firstRows + combinedRow, - q4kRowDot( - firstWeights, - combinedRow, - activations, - batch * cols, - activationScales, - batch * blocks, - activationSums, - batch * sumsPerBatch, - blocks)); - } else if (combinedRow < firstAndSecondRows) { - int row = combinedRow - firstRows; - secondOutput.set( - batch * secondRows + row, - q4kRowDot( - secondWeights, - row, - activations, - batch * cols, - activationScales, - batch * blocks, - activationSums, - batch * sumsPerBatch, - blocks)); - } else { - int row = combinedRow - firstAndSecondRows; - thirdOutput.set( - batch * thirdRows + row, - q6kRowDot( - thirdWeights, - row, - activations, - batch * cols, - activationScales, - batch * blocks, - blocks)); - } - } - } - - private static float q4kRowDot( - ByteArray weights, - int row, - ByteArray activations, - int activationOffset, - FloatArray activationScales, - int scaleOffset, - IntArray activationSums, - int sumOffset, - int blocks) { - int rowOffset = row * blocks * Q4_K_BLOCK_BYTES; - float sum = 0.0f; - for (int block = 0; block < blocks; block++) { - int blockOffset = rowOffset + block * Q4_K_BLOCK_BYTES; - float activationScale = activationScales.get(scaleOffset + block); - float delta = weights.getHalfFloat(blockOffset).getFloat32() * activationScale; - float deltaMin = weights.getHalfFloat(blockOffset + 2).getFloat32() * activationScale; - int scalesOffset = blockOffset + Q4_K_SCALES_OFFSET; - int quantsOffset = blockOffset + Q4_K_QUANTS_OFFSET; - int blockActivationOffset = activationOffset + block * SUPER_BLOCK_VALUES; - int blockSumOffset = sumOffset + block * SUMS_PER_SUPER_BLOCK; - int quantizedSum = 0; - int minimumSum = 0; - for (int group = 0; group < Q4_K_GROUPS; group++) { - int packedOffset = quantsOffset + (group >>> 1) * Q4_K_GROUP_VALUES; - int shift = (group & 1) * 4; - int groupActivationOffset = blockActivationOffset + group * Q4_K_GROUP_VALUES; - int groupDot = 0; - for (int index = 0; index < Q4_K_GROUP_VALUES; index++) { - int packed = weights.get(packedOffset + index) & 0xFF; - int quant = (packed >>> shift) & 0x0F; - groupDot += quant * activations.get(groupActivationOffset + index); - } - quantizedSum += groupScale(weights, scalesOffset, group) * groupDot; - int groupSumOffset = blockSumOffset + group * 2; - minimumSum += - groupMinimum(weights, scalesOffset, group) - * (activationSums.get(groupSumOffset) + activationSums.get(groupSumOffset + 1)); - } - sum = sum + delta * quantizedSum; - sum = sum - deltaMin * minimumSum; - } - return sum; - } - - private static float q6kRowDot( - ByteArray weights, - int row, - ByteArray activations, - int activationOffset, - FloatArray activationScales, - int scaleOffset, - int blocks) { - int rowOffset = row * blocks * Q6_K_BLOCK_BYTES; - float sum = 0.0f; - for (int block = 0; block < blocks; block++) { - int blockOffset = rowOffset + block * Q6_K_BLOCK_BYTES; - float delta = - weights.getHalfFloat(blockOffset + Q6_K_DELTA_OFFSET).getFloat32() - * activationScales.get(scaleOffset + block); - int blockActivationOffset = activationOffset + block * SUPER_BLOCK_VALUES; - int blockSum = 0; - for (int half = 0; half < 2; half++) { - int lowBase = blockOffset + half * 64; - int highBase = blockOffset + Q6_K_QL_BYTES + half * 32; - int scaleBase = blockOffset + Q6_K_SCALE_OFFSET + half * 8; - int quantBase = blockActivationOffset + half * 128; - for (int index = 0; index < 32; index++) { - int scaleIndex = index >>> 4; - int lowFirst = weights.get(lowBase + index) & 0xFF; - int lowSecond = weights.get(lowBase + 32 + index) & 0xFF; - int high = weights.get(highBase + index) & 0xFF; - int first = ((lowFirst & 0x0F) | ((high & 0x03) << 4)) - 32; - int second = ((lowSecond & 0x0F) | (((high >>> 2) & 0x03) << 4)) - 32; - int third = ((lowFirst >>> 4) | (((high >>> 4) & 0x03) << 4)) - 32; - int fourth = ((lowSecond >>> 4) | (((high >>> 6) & 0x03) << 4)) - 32; - blockSum += - weights.get(scaleBase + scaleIndex) * first * activations.get(quantBase + index) - + weights.get(scaleBase + scaleIndex + 2) - * second - * activations.get(quantBase + index + 32) - + weights.get(scaleBase + scaleIndex + 4) - * third - * activations.get(quantBase + index + 64) - + weights.get(scaleBase + scaleIndex + 6) - * fourth - * activations.get(quantBase + index + 96); - } - } - sum = sum + delta * blockSum; - } - return sum; - } - - /** Decodes one of the eight six-bit group scales packed into a Q4_K super-block header. */ - private static int groupScale(ByteArray weights, int scalesOffset, int group) { - if (group < 4) { - return weights.get(scalesOffset + group) & 0x3F; - } - int low = weights.get(scalesOffset + group + 4) & 0x0F; - int high = (weights.get(scalesOffset + group - 4) & 0xFF) >>> 6; - return low | (high << 4); - } - - /** Decodes one of the eight six-bit group minima packed into a Q4_K super-block header. */ - private static int groupMinimum(ByteArray weights, int scalesOffset, int group) { - if (group < 4) { - return weights.get(scalesOffset + group + 4) & 0x3F; - } - int low = (weights.get(scalesOffset + group + 4) & 0xFF) >>> 4; - int high = (weights.get(scalesOffset + group) & 0xFF) >>> 6; - return low | (high << 4); - } - - /** - * Reproduces the production Q8_K activation preparation. - * - *

Unlike Q8_0, the Q8_K scale is not rounded through fp16 and the block is scaled by the - * signed extremum rather than the absolute maximum. The per-16-value sums are what Q4_K's minimum - * correction consumes; they always fit in a {@code short} on the CPU side, so widening them to - * {@code int} for the device is value-preserving. - */ - public static void quantize( - float[] input, byte[] activations, float[] scales, int[] sums, int batchSize, int cols) { - int blocks = cols / SUPER_BLOCK_VALUES; - for (int batch = 0; batch < batchSize; batch++) { - int rowOffset = batch * cols; - int rowScaleOffset = batch * blocks; - int rowSumOffset = batch * (cols / SUM_BLOCK_VALUES); - for (int block = 0; block < blocks; block++) { - int offset = rowOffset + block * SUPER_BLOCK_VALUES; - float extremum = 0.0f; - float absoluteMax = 0.0f; - for (int index = 0; index < SUPER_BLOCK_VALUES; index++) { - float value = input[offset + index]; - float absolute = Math.abs(value); - if (absolute > absoluteMax) { - absoluteMax = absolute; - extremum = value; - } - } - int blockSumOffset = rowSumOffset + block * SUMS_PER_SUPER_BLOCK; - if (absoluteMax == 0.0f) { - for (int index = 0; index < SUPER_BLOCK_VALUES; index++) { - activations[offset + index] = 0; - } - for (int index = 0; index < SUMS_PER_SUPER_BLOCK; index++) { - sums[blockSumOffset + index] = 0; - } - scales[rowScaleOffset + block] = 0.0f; - continue; - } - float inverseScale = -127.0f / extremum; - int partialSum = 0; - for (int index = 0; index < SUPER_BLOCK_VALUES; index++) { - int quant = ggmlNearestInt(inverseScale * input[offset + index]); - byte stored = (byte) Math.min(127, quant); - activations[offset + index] = stored; - partialSum += stored; - if ((index + 1) % SUM_BLOCK_VALUES == 0) { - sums[blockSumOffset + index / SUM_BLOCK_VALUES] = partialSum; - partialSum = 0; - } - } - scales[rowScaleOffset + block] = 1.0f / inverseScale; - } - } - } - - /** Validates one Q4_K projection's device storage before a task graph is built. */ - public static void validateQ4K( - ByteArray weights, - ByteArray activations, - FloatArray activationScales, - IntArray activationSums, - FloatArray output, - int batchSize, - int rows, - int cols) { - validateShape(batchSize, rows, cols); - validateWeights(weights, rows, cols, Q4_K_BLOCK_BYTES); - validateActivations(activations, activationScales, batchSize, cols); - if (activationSums == null - || activationSums.getSize() != Math.multiplyExact(batchSize, cols / SUM_BLOCK_VALUES)) { - throw new IllegalArgumentException("activation sums do not match the projection shape"); - } - validateOutput(output, batchSize, rows); - } - - /** Validates one Q6_K projection's device storage before a task graph is built. */ - public static void validateQ6K( - ByteArray weights, - ByteArray activations, - FloatArray activationScales, - FloatArray output, - int batchSize, - int rows, - int cols) { - validateShape(batchSize, rows, cols); - validateWeights(weights, rows, cols, Q6_K_BLOCK_BYTES); - validateActivations(activations, activationScales, batchSize, cols); - validateOutput(output, batchSize, rows); - } - - private static void validateShape(int batchSize, int rows, int cols) { - if (batchSize < 1) { - throw new IllegalArgumentException("batchSize must be positive"); - } - if (rows < 1) { - throw new IllegalArgumentException("rows must be positive"); - } - if (cols < SUPER_BLOCK_VALUES || cols % SUPER_BLOCK_VALUES != 0) { - throw new IllegalArgumentException("cols must be a positive multiple of 256"); - } - } - - private static void validateWeights(ByteArray weights, int rows, int cols, int blockBytes) { - if (weights == null) { - throw new IllegalArgumentException("weights do not match the projection shape"); - } - int blocks = cols / SUPER_BLOCK_VALUES; - long expected = (long) rows * blocks * blockBytes; - if (expected > Integer.MAX_VALUE || weights.getSize() != (int) expected) { - throw new IllegalArgumentException("weights do not match the projection shape"); - } - } - - private static void validateActivations( - ByteArray activations, FloatArray activationScales, int batchSize, int cols) { - if (activations == null || activations.getSize() != Math.multiplyExact(batchSize, cols)) { - throw new IllegalArgumentException("activations do not match the projection shape"); - } - if (activationScales == null - || activationScales.getSize() != Math.multiplyExact(batchSize, cols / SUPER_BLOCK_VALUES)) { - throw new IllegalArgumentException("activation scales do not match the projection shape"); - } - } - - private static void validateOutput(FloatArray output, int batchSize, int rows) { - if (output == null || output.getSize() != Math.multiplyExact(batchSize, rows)) { - throw new IllegalArgumentException("output does not match the projection shape"); - } - } - - private static int ggmlNearestInt(float value) { - int bits = Float.floatToRawIntBits(value + 12_582_912.0f); - return (bits & 0x007F_FFFF) - 0x0040_0000; - } -} diff --git a/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/PlanShapeStrategy.java b/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/PlanShapeStrategy.java deleted file mode 100644 index 9fa2a34f0..000000000 --- a/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/PlanShapeStrategy.java +++ /dev/null @@ -1,108 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.backend.tornado; - -/** - * How retained TornadoVM execution plans hold model weights on the device. - * - *

Only {@link #PER_SHAPE_WHOLE_MODEL} is implemented today. The other constants exist so the - * capacity gate can say precisely which unimplemented plan shape a model would have needed, instead - * of reporting an anonymous shortfall. - */ -public enum PlanShapeStrategy { - - /** - * One {@code TaskGraph} per (weight tensor, batch shape), each owning its own device weight - * buffer and its own host copy. - * - *

This is what {@code TornadoGgufBatchedMatrixKernel} builds: the plan cache is keyed on the - * weight address plus the execution batch size, and the constructor calls {@code - * ByteArray.fromSegment(weights)}, which copies the tensor into a fresh off-heap array before the - * task graph pins it with {@code DataTransferMode.FIRST_EXECUTION}. With decode acceleration on, - * every weight is therefore held twice on the device and twice on the host. - */ - PER_SHAPE_WHOLE_MODEL(true, null), - - /** - * One device weight buffer per tensor, shared by every batch shape that reads it. - * - *

Halves both device and host weight residency when prefill and decode plans are retained - * together. - */ - SHARED_WEIGHT_UPLOAD( - false, - "TornadoGgufBatchedMatrixKernel allocates a new ByteArray per (weight, batch shape) plan;" - + " sharing one device buffer across shapes requires hoisting the weight array out of" - + " the plan constructor and keying it on the weight address alone"), - - /** - * A bounded resident working set: weights stay on the host and only the tensors needed by the - * current layer or routed experts are uploaded, under an LRU bound. - * - *

Device capacity stops being the binding constraint and interconnect bandwidth becomes it, so - * this strategy is never selected by the capacity gate; it is named only so an ineligibility - * message can point at it. - */ - BOUNDED_RESIDENT_WORKING_SET( - false, - "no eviction path exists: plans are held in unbounded LinkedHashMaps and only released on" - + " kernel close, and per-token re-upload is bounded by PCIe bandwidth rather than by" - + " device capacity"); - - private final boolean implemented; - private final String limitation; - - PlanShapeStrategy(boolean implemented, String limitation) { - this.implemented = implemented; - this.limitation = limitation; - } - - /** Whether the shipped kernel can actually build plans in this shape. */ - public boolean implemented() { - return implemented; - } - - /** What stops this strategy from being available, or {@code null} when it is implemented. */ - public String limitation() { - return limitation; - } - - /** Whether device capacity for this strategy can be computed from a request. */ - public boolean capacityComputable() { - return this != BOUNDED_RESIDENT_WORKING_SET; - } - - /** Device bytes of model weights retained under this strategy. */ - long deviceWeightBytes(DeviceMemoryRequest request) { - return switch (this) { - case PER_SHAPE_WHOLE_MODEL -> - Math.multiplyExact(request.weightBytes(), request.retainedShapes()); - case SHARED_WEIGHT_UPLOAD -> request.weightBytes(); - case BOUNDED_RESIDENT_WORKING_SET -> - throw new IllegalStateException( - "BOUNDED_RESIDENT_WORKING_SET has no capacity-derived weight budget"); - }; - } - - /** - * Host bytes of off-heap weight copies retained under this strategy. - * - *

{@code ByteArray.fromSegment} copies; the mapped GGUF stays mapped alongside these copies. - */ - long hostWeightCopyBytes(DeviceMemoryRequest request) { - return deviceWeightBytes(request); - } -} diff --git a/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/Q4ProjectionKernel.java b/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/Q4ProjectionKernel.java deleted file mode 100644 index 22d379862..000000000 --- a/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/Q4ProjectionKernel.java +++ /dev/null @@ -1,348 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.backend.tornado; - -import uk.ac.manchester.tornado.api.annotations.Parallel; -import uk.ac.manchester.tornado.api.types.arrays.ByteArray; -import uk.ac.manchester.tornado.api.types.arrays.FloatArray; - -/** Java-authored TornadoVM kernel for production-compatible Q4_0 by Q8_0 projections. */ -public final class Q4ProjectionKernel { - private static final int BLOCK_VALUES = 32; - private static final int BLOCK_BYTES = 18; - - private Q4ProjectionKernel() {} - - /** Runs one work item per batch/output-row pair over host-prepared Q8_0 activations. */ - public static void multiply( - ByteArray weights, - ByteArray activations, - FloatArray activationScales, - FloatArray output, - int batchSize, - int rows, - int cols) { - int blocks = cols / BLOCK_VALUES; - for (@Parallel int outputIndex = 0; outputIndex < batchSize * rows; outputIndex++) { - int batch = outputIndex / rows; - int row = outputIndex - batch * rows; - int activationOffset = batch * cols; - int activationScaleOffset = batch * blocks; - output.set( - outputIndex, - rowDotLow( - weights, - row, - activations, - activationOffset, - activationScales, - activationScaleOffset, - blocks) - + rowDotHigh( - weights, - row, - activations, - activationOffset, - activationScales, - activationScaleOffset, - blocks)); - } - } - - /** Computes two projections in one dispatch over one prepared activation. */ - public static void multiplyDual( - ByteArray firstWeights, - int firstRows, - ByteArray secondWeights, - int secondRows, - ByteArray activations, - FloatArray activationScales, - FloatArray firstOutput, - FloatArray secondOutput, - int batchSize, - int cols) { - int blocks = cols / BLOCK_VALUES; - int combinedRows = firstRows + secondRows; - for (@Parallel int outputIndex = 0; outputIndex < batchSize * combinedRows; outputIndex++) { - int batch = outputIndex / combinedRows; - int combinedRow = outputIndex - batch * combinedRows; - int activationOffset = batch * cols; - int activationScaleOffset = batch * blocks; - if (combinedRow < firstRows) { - firstOutput.set( - batch * firstRows + combinedRow, - rowDotLow( - firstWeights, - combinedRow, - activations, - activationOffset, - activationScales, - activationScaleOffset, - blocks) - + rowDotHigh( - firstWeights, - combinedRow, - activations, - activationOffset, - activationScales, - activationScaleOffset, - blocks)); - } else { - int row = combinedRow - firstRows; - secondOutput.set( - batch * secondRows + row, - rowDotLow( - secondWeights, - row, - activations, - activationOffset, - activationScales, - activationScaleOffset, - blocks) - + rowDotHigh( - secondWeights, - row, - activations, - activationOffset, - activationScales, - activationScaleOffset, - blocks)); - } - } - } - - /** Computes three projections in one dispatch over one prepared activation. */ - public static void multiplyTriple( - ByteArray firstWeights, - int firstRows, - ByteArray secondWeights, - int secondRows, - ByteArray thirdWeights, - int thirdRows, - ByteArray activations, - FloatArray activationScales, - FloatArray firstOutput, - FloatArray secondOutput, - FloatArray thirdOutput, - int batchSize, - int cols) { - int blocks = cols / BLOCK_VALUES; - int firstAndSecondRows = firstRows + secondRows; - int combinedRows = firstAndSecondRows + thirdRows; - for (@Parallel int outputIndex = 0; outputIndex < batchSize * combinedRows; outputIndex++) { - int batch = outputIndex / combinedRows; - int combinedRow = outputIndex - batch * combinedRows; - int activationOffset = batch * cols; - int activationScaleOffset = batch * blocks; - if (combinedRow < firstRows) { - firstOutput.set( - batch * firstRows + combinedRow, - rowDotLow( - firstWeights, - combinedRow, - activations, - activationOffset, - activationScales, - activationScaleOffset, - blocks) - + rowDotHigh( - firstWeights, - combinedRow, - activations, - activationOffset, - activationScales, - activationScaleOffset, - blocks)); - } else if (combinedRow < firstAndSecondRows) { - int row = combinedRow - firstRows; - secondOutput.set( - batch * secondRows + row, - rowDotLow( - secondWeights, - row, - activations, - activationOffset, - activationScales, - activationScaleOffset, - blocks) - + rowDotHigh( - secondWeights, - row, - activations, - activationOffset, - activationScales, - activationScaleOffset, - blocks)); - } else { - int row = combinedRow - firstAndSecondRows; - thirdOutput.set( - batch * thirdRows + row, - rowDotLow( - thirdWeights, - row, - activations, - activationOffset, - activationScales, - activationScaleOffset, - blocks) - + rowDotHigh( - thirdWeights, - row, - activations, - activationOffset, - activationScales, - activationScaleOffset, - blocks)); - } - } - } - - private static float rowDotLow( - ByteArray weights, - int row, - ByteArray activations, - int activationOffset, - FloatArray activationScales, - int activationScaleOffset, - int blocks) { - int rowBlockOffset = row * blocks; - float sum = 0.0f; - for (int block = 0; block < blocks; block++) { - int weightOffset = (rowBlockOffset + block) * BLOCK_BYTES; - int quantOffset = activationOffset + block * BLOCK_VALUES; - int integerSum = pairProductSumLow(weights, weightOffset, activations, quantOffset); - float scale = - weights.getHalfFloat(weightOffset).getFloat32() - * activationScales.get(activationScaleOffset + block); - sum += scale * integerSum; - } - return sum; - } - - private static float rowDotHigh( - ByteArray weights, - int row, - ByteArray activations, - int activationOffset, - FloatArray activationScales, - int activationScaleOffset, - int blocks) { - int rowBlockOffset = row * blocks; - float sum = 0.0f; - for (int block = 0; block < blocks; block++) { - int weightOffset = (rowBlockOffset + block) * BLOCK_BYTES; - int quantOffset = activationOffset + block * BLOCK_VALUES; - int integerSum = pairProductSumHigh(weights, weightOffset, activations, quantOffset); - float scale = - weights.getHalfFloat(weightOffset).getFloat32() - * activationScales.get(activationScaleOffset + block); - sum += scale * integerSum; - } - return sum; - } - - private static int pairProductSumLow( - ByteArray weights, int weightOffset, ByteArray activations, int activationOffset) { - return pairProduct(weights, weightOffset, activations, activationOffset, 0) - + pairProduct(weights, weightOffset, activations, activationOffset, 1) - + pairProduct(weights, weightOffset, activations, activationOffset, 2) - + pairProduct(weights, weightOffset, activations, activationOffset, 3) - + pairProduct(weights, weightOffset, activations, activationOffset, 4) - + pairProduct(weights, weightOffset, activations, activationOffset, 5) - + pairProduct(weights, weightOffset, activations, activationOffset, 6) - + pairProduct(weights, weightOffset, activations, activationOffset, 7); - } - - private static int pairProductSumHigh( - ByteArray weights, int weightOffset, ByteArray activations, int activationOffset) { - return pairProduct(weights, weightOffset, activations, activationOffset, 8) - + pairProduct(weights, weightOffset, activations, activationOffset, 9) - + pairProduct(weights, weightOffset, activations, activationOffset, 10) - + pairProduct(weights, weightOffset, activations, activationOffset, 11) - + pairProduct(weights, weightOffset, activations, activationOffset, 12) - + pairProduct(weights, weightOffset, activations, activationOffset, 13) - + pairProduct(weights, weightOffset, activations, activationOffset, 14) - + pairProduct(weights, weightOffset, activations, activationOffset, 15); - } - - private static int pairProduct( - ByteArray weights, int weightOffset, ByteArray activations, int activationOffset, int index) { - int packed = weights.get(weightOffset + 2 + index) & 0xFF; - int low = (packed & 0x0F) - 8; - int high = ((packed >>> 4) & 0x0F) - 8; - return low * activations.get(activationOffset + index) - + high * activations.get(activationOffset + index + 16); - } - - /** Reproduces the production Q8_0 activation preparation, including FP16 scale rounding. */ - public static void quantize( - float[] input, byte[] activations, float[] scales, int batchSize, int cols) { - int blocks = cols / BLOCK_VALUES; - for (int batch = 0; batch < batchSize; batch++) { - for (int block = 0; block < blocks; block++) { - int offset = batch * cols + block * BLOCK_VALUES; - float absoluteMax = 0.0f; - for (int index = 0; index < BLOCK_VALUES; index++) { - absoluteMax = Math.max(absoluteMax, Math.abs(input[offset + index])); - } - float scale = absoluteMax / 127.0f; - float inverseScale = absoluteMax == 0.0f ? 0.0f : 127.0f / absoluteMax; - scales[batch * blocks + block] = Float.float16ToFloat(Float.floatToFloat16(scale)); - for (int index = 0; index < BLOCK_VALUES; index++) { - activations[offset + index] = (byte) ggmlNearestInt(input[offset + index] * inverseScale); - } - } - } - } - - public static void validate( - ByteArray weights, - ByteArray activations, - FloatArray activationScales, - FloatArray output, - int batchSize, - int rows, - int cols) { - if (batchSize < 1) { - throw new IllegalArgumentException("batchSize must be positive"); - } - if (rows < 1) { - throw new IllegalArgumentException("rows must be positive"); - } - if (cols < BLOCK_VALUES || cols % BLOCK_VALUES != 0) { - throw new IllegalArgumentException("cols must be a positive multiple of 32"); - } - int blocks = cols / BLOCK_VALUES; - int expectedWeightBytes = Math.multiplyExact(Math.multiplyExact(rows, blocks), BLOCK_BYTES); - if (weights.getSize() != expectedWeightBytes) { - throw new IllegalArgumentException("weights do not match the projection shape"); - } - if (activations.getSize() != Math.multiplyExact(batchSize, cols)) { - throw new IllegalArgumentException("activations do not match the projection shape"); - } - if (activationScales.getSize() != Math.multiplyExact(batchSize, blocks)) { - throw new IllegalArgumentException("activation scales do not match the projection shape"); - } - if (output.getSize() != Math.multiplyExact(batchSize, rows)) { - throw new IllegalArgumentException("output does not match the projection shape"); - } - } - - private static int ggmlNearestInt(float value) { - int bits = Float.floatToRawIntBits(value + 12_582_912.0f); - return (bits & 0x007F_FFFF) - 0x0040_0000; - } -} diff --git a/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/TornadoBackend.java b/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/TornadoBackend.java deleted file mode 100644 index 25039add0..000000000 --- a/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/TornadoBackend.java +++ /dev/null @@ -1,172 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.backend.tornado; - -import com.integrallis.models.api.BackendConfiguration; -import com.integrallis.models.backend.purejava.PureJavaBackend; -import com.integrallis.models.backend.purejava.spi.GgufBatchedMatrixKernel; -import java.io.IOException; -import java.io.UncheckedIOException; -import java.nio.file.Files; -import java.nio.file.Path; -import java.time.Duration; -import java.util.Arrays; -import java.util.List; -import java.util.Objects; - -/** Loads the qualified Java/Tornado Q4 backend when possible and otherwise uses the Vector API. */ -public final class TornadoBackend { - private static final System.Logger LOGGER = System.getLogger(TornadoBackend.class.getName()); - - private TornadoBackend() {} - - /** Opens a model with automatic device selection, eager readiness, and safe CPU fallback. */ - public static TornadoBackendRuntime open(Path modelPath) { - return open(modelPath, BackendConfiguration.empty(), TornadoBackendOptions.defaults()); - } - - /** Opens a model using explicit accelerator readiness and fallback controls. */ - public static TornadoBackendRuntime open(Path modelPath, TornadoBackendOptions options) { - return open(modelPath, BackendConfiguration.empty(), options); - } - - /** Opens a model using its qualified backend configuration and accelerator controls. */ - public static TornadoBackendRuntime open( - Path modelPath, BackendConfiguration backendConfiguration, TornadoBackendOptions options) { - Objects.requireNonNull(modelPath, "modelPath"); - Objects.requireNonNull(backendConfiguration, "backendConfiguration"); - Objects.requireNonNull(options, "options"); - Path model = modelPath.toAbsolutePath().normalize(); - if (!Files.isRegularFile(model)) { - throw new IllegalArgumentException("modelPath must be a regular model file: " + model); - } - long modelBytes = size(model); - List devices; - try { - devices = TornadoRuntimeDevices.discover(); - } catch (LinkageError | RuntimeException failure) { - return fallback( - model, - backendConfiguration, - options, - "TornadoVM runtime unavailable: " + failure.getClass().getSimpleName(), - 0); - } - AcceleratorEligibility.Decision decision = - AcceleratorEligibility.select( - devices, - DeviceMemoryRequest.ofModelFile( - modelLabel(model), modelBytes, options.accelerateDecode())); - if (!decision.eligible()) { - return fallback( - model, backendConfiguration, options, decision.reason(), decision.requiredBytes()); - } - - TornadoGgufBatchedMatrixKernel kernel = - new TornadoGgufBatchedMatrixKernel( - options.executionBatchSize(), options.accelerateDecode()); - PureJavaBackend backend = null; - try { - backend = PureJavaBackend.load(model, backendConfiguration, kernel); - Duration readiness = options.eagerReadiness() ? prepare(backend, options) : Duration.ZERO; - if (options.eagerReadiness() && kernel.calls() == 0) { - backend.close(); - return fallback( - model, - backendConfiguration, - options, - "model has no eligible Q4_0 or K-quant projections", - decision.requiredBytes()); - } - String device = decision.device().name(); - LOGGER.log( - System.Logger.Level.INFO, - "Models accelerator selected {0}; readiness={1} ms plans={2} routed={3}", - device, - readiness.toMillis(), - kernel.projectionPlanCount(), - kernel.routedProjectionsByFormat()); - return new TornadoBackendRuntime( - backend, - new TornadoBackendStatus(true, device, "eligible", decision.requiredBytes(), readiness), - kernel); - } catch (LinkageError | RuntimeException failure) { - if (backend != null) { - try { - backend.close(); - } catch (RuntimeException closeFailure) { - failure.addSuppressed(closeFailure); - } - } - if (options.requireAccelerator()) { - throw new IllegalStateException("qualified accelerator initialization failed", failure); - } - return fallback( - model, - backendConfiguration, - options, - "accelerator initialization failed: " + failure.getClass().getSimpleName(), - decision.requiredBytes()); - } - } - - private static Duration prepare(PureJavaBackend backend, TornadoBackendOptions options) { - long started = System.nanoTime(); - int token = readinessToken(backend); - int[] tokens = new int[options.executionBatchSize()]; - Arrays.fill(tokens, token); - backend.prefill(tokens, 0); - if (options.accelerateDecode()) { - backend.forward(token, tokens.length); - } - backend.reset(); - return Duration.ofNanos(System.nanoTime() - started); - } - - private static int readinessToken(PureJavaBackend backend) { - int token = backend.tokenizer().bosToken(); - return token >= 0 && token < backend.metadata().vocabSize() ? token : 0; - } - - private static TornadoBackendRuntime fallback( - Path model, - BackendConfiguration backendConfiguration, - TornadoBackendOptions options, - String reason, - long requiredBytes) { - if (options.requireAccelerator()) { - throw new IllegalStateException("accelerator required but unavailable: " + reason); - } - LOGGER.log( - System.Logger.Level.INFO, "Models accelerator unavailable; using Vector API: {0}", reason); - return new TornadoBackendRuntime( - PureJavaBackend.load(model, backendConfiguration, GgufBatchedMatrixKernel.none()), - new TornadoBackendStatus(false, "Vector API", reason, requiredBytes, Duration.ZERO)); - } - - private static String modelLabel(Path model) { - Path fileName = model.getFileName(); - return fileName == null ? model.toString() : fileName.toString(); - } - - private static long size(Path model) { - try { - return Files.size(model); - } catch (IOException exception) { - throw new UncheckedIOException("could not read model size: " + model, exception); - } - } -} diff --git a/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/TornadoBackendOptions.java b/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/TornadoBackendOptions.java deleted file mode 100644 index eab6fbd5c..000000000 --- a/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/TornadoBackendOptions.java +++ /dev/null @@ -1,35 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.backend.tornado; - -/** Immutable controls for optional TornadoVM device selection and readiness. */ -public record TornadoBackendOptions( - boolean accelerateDecode, - boolean eagerReadiness, - boolean requireAccelerator, - int executionBatchSize) { - - public TornadoBackendOptions { - if (executionBatchSize < 4) { - throw new IllegalArgumentException("executionBatchSize must be at least 4"); - } - } - - /** Selects qualified hardware automatically and falls back to the Java Vector API. */ - public static TornadoBackendOptions defaults() { - return new TornadoBackendOptions(true, true, false, 32); - } -} diff --git a/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/TornadoBackendRuntime.java b/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/TornadoBackendRuntime.java deleted file mode 100644 index b0bc7eb48..000000000 --- a/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/TornadoBackendRuntime.java +++ /dev/null @@ -1,81 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.backend.tornado; - -import com.integrallis.models.backend.purejava.PureJavaBackend; -import java.util.Map; -import java.util.Objects; - -/** Owns the automatically selected accelerated or Vector API backend. */ -public final class TornadoBackendRuntime implements AutoCloseable { - private PureJavaBackend backend; - private final TornadoBackendStatus status; - private final TornadoGgufBatchedMatrixKernel kernel; - - TornadoBackendRuntime(PureJavaBackend backend, TornadoBackendStatus status) { - this(backend, status, null); - } - - TornadoBackendRuntime( - PureJavaBackend backend, TornadoBackendStatus status, TornadoGgufBatchedMatrixKernel kernel) { - this.backend = Objects.requireNonNull(backend, "backend"); - this.status = Objects.requireNonNull(status, "status"); - this.kernel = kernel; - } - - /** Returns the loaded backend used by the ordinary Models generation pipeline. */ - public PureJavaBackend backend() { - if (backend == null) { - throw new IllegalStateException("backend ownership was transferred"); - } - return backend; - } - - PureJavaBackend detachBackend() { - PureJavaBackend detached = backend(); - backend = null; - return detached; - } - - /** Returns the device-selection, fallback, and readiness outcome. */ - public TornadoBackendStatus status() { - return status; - } - - /** - * Returns the projections routed to the device so far, keyed by GGUF weight format. - * - *

Empty when the load fell back to the Vector API. A grouped dispatch counts once per matrix, - * so a mixed-format model's counts show whether each of its formats reached the device rather - * than only the majority one. - */ - public Map routedProjectionsByFormat() { - return kernel == null ? Map.of() : kernel.routedProjectionsByFormat(); - } - - /** Number of distinct compiled device execution plans, or zero when not accelerated. */ - public int projectionPlanCount() { - return kernel == null ? 0 : kernel.projectionPlanCount(); - } - - @Override - public void close() { - if (backend != null) { - backend.close(); - backend = null; - } - } -} diff --git a/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/TornadoBackendStatus.java b/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/TornadoBackendStatus.java deleted file mode 100644 index 6b9b260a0..000000000 --- a/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/TornadoBackendStatus.java +++ /dev/null @@ -1,48 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.backend.tornado; - -import java.time.Duration; -import java.util.Objects; - -/** Immutable outcome of automatic accelerator selection and readiness. */ -public record TornadoBackendStatus( - boolean accelerated, - String device, - String reason, - long requiredDeviceBytes, - Duration readinessTime) { - - public TornadoBackendStatus { - device = requireText(device, "device"); - reason = requireText(reason, "reason"); - if (requiredDeviceBytes < 0) { - throw new IllegalArgumentException("requiredDeviceBytes must not be negative"); - } - readinessTime = Objects.requireNonNull(readinessTime, "readinessTime"); - if (readinessTime.isNegative()) { - throw new IllegalArgumentException("readinessTime must not be negative"); - } - } - - private static String requireText(String value, String name) { - Objects.requireNonNull(value, name); - if (value.isBlank()) { - throw new IllegalArgumentException(name + " must not be blank"); - } - return value.strip(); - } -} diff --git a/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/TornadoGgufBatchedMatrixKernel.java b/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/TornadoGgufBatchedMatrixKernel.java deleted file mode 100644 index d4d27235f..000000000 --- a/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/TornadoGgufBatchedMatrixKernel.java +++ /dev/null @@ -1,1091 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.backend.tornado; - -import com.integrallis.models.backend.purejava.gguf.GgufTensorType; -import com.integrallis.models.backend.purejava.plan.PureJavaPlanConfiguration; -import com.integrallis.models.backend.purejava.spi.GgufBatchedMatrixKernel; -import java.lang.foreign.MemorySegment; -import java.util.ArrayList; -import java.util.Arrays; -import java.util.Collections; -import java.util.EnumMap; -import java.util.LinkedHashMap; -import java.util.List; -import java.util.Map; -import java.util.Objects; -import uk.ac.manchester.tornado.api.TaskGraph; -import uk.ac.manchester.tornado.api.TornadoExecutionPlan; -import uk.ac.manchester.tornado.api.enums.DataTransferMode; -import uk.ac.manchester.tornado.api.types.arrays.ByteArray; -import uk.ac.manchester.tornado.api.types.arrays.FloatArray; -import uk.ac.manchester.tornado.api.types.arrays.IntArray; - -/** - * Q4_0 and K-quant prefill and optional decode projections backed by Java-authored TornadoVM - * kernels. - * - *

Two activation families are served. Q4_0 weights consume the Q8_0 activations prepared by - * {@link Q4ProjectionKernel}; Q4_K and Q6_K weights consume the Q8_K activations prepared by {@link - * KQuantProjectionKernel}. A grouped dispatch shares one activation preparation, so the two - * families are never mixed inside one dual or triple projection even though a Q4_K_M model contains - * both. - */ -public final class TornadoGgufBatchedMatrixKernel implements GgufBatchedMatrixKernel { - private static final int MINIMUM_BATCH = 4; - private static final int DEFAULT_EXECUTION_BATCH_SIZE = 32; - private static final long MINIMUM_MATRIX_VALUES = 1_048_576L; - - /** - * Largest weight tensor a single execution plan can hold on the device. - * - *

TornadoVM's {@code ByteArray} is indexed with an {@code int} and carries a small header, so - * a mapped tensor at or above two gibibytes cannot be addressed at all. Rejecting it here keeps - * the projection on the Vector API rather than failing the load. - */ - private static final long MAX_DEVICE_TENSOR_BYTES = Integer.MAX_VALUE - 1024L; - - private static final Map PLAN_RECOMMENDATIONS = - Map.of( - PureJavaPlanConfiguration.GROUPED_PROJECTIONS_PROPERTY, - "true", - PureJavaPlanConfiguration.MIXED_K_PROJECTIONS_PROPERTY, - "true", - PureJavaPlanConfiguration.STAGED_QUANTIZED_FFN_PROPERTY, - "false", - PureJavaPlanConfiguration.STAGED_QUANTIZED_LAYER_PROPERTY, - "false"); - - private final Map plans = new LinkedHashMap<>(); - private final Map dualPlans = new LinkedHashMap<>(); - private final Map triplePlans = new LinkedHashMap<>(); - private final Map kQuantPlans = new LinkedHashMap<>(); - private final Map routedProjections = new EnumMap<>(GgufTensorType.class); - private final int executionBatchSize; - private final boolean accelerateDecode; - private int planSequence; - private long calls; - private long totalNanos; - private boolean closed; - - public TornadoGgufBatchedMatrixKernel() { - this(DEFAULT_EXECUTION_BATCH_SIZE, false); - } - - public TornadoGgufBatchedMatrixKernel(int executionBatchSize) { - this(executionBatchSize, false); - } - - public TornadoGgufBatchedMatrixKernel(int executionBatchSize, boolean accelerateDecode) { - if (executionBatchSize < MINIMUM_BATCH) { - throw new IllegalArgumentException("executionBatchSize must be at least " + MINIMUM_BATCH); - } - this.executionBatchSize = executionBatchSize; - this.accelerateDecode = accelerateDecode; - } - - /** Fixed device batch shape used to make compiled plans reusable across prompt lengths. */ - public int executionBatchSize() { - return executionBatchSize; - } - - /** Whether this provider creates separate single-token projection plans for decode. */ - public boolean acceleratesDecode() { - return accelerateDecode; - } - - int executionBatchSizeFor(int actualBatchSize) { - return accelerateDecode && actualBatchSize == 1 ? 1 : executionBatchSize; - } - - @Override - public String implementation() { - return accelerateDecode ? "tornadovm-java-q4-prefill-decode" : "tornadovm-java-q4-prefill"; - } - - @Override - public Map planRecommendations() { - return PLAN_RECOMMENDATIONS; - } - - @Override - public boolean supports(GgufTensorType type) { - return type == GgufTensorType.Q4_0 - || type == GgufTensorType.Q4_K - || type == GgufTensorType.Q6_K; - } - - @Override - public boolean isEligible(GgufTensorType type, int batchSize, int rows, int cols) { - return supports(type) - && eligibleBatchSize(batchSize) - && batchSize <= executionBatchSize - && rows > 0 - && cols > 0 - && cols % type.blockSize() == 0 - && deviceAddressable(type, rows, cols) - && (long) rows * cols >= MINIMUM_MATRIX_VALUES; - } - - @Override - public boolean supportsDual(GgufTensorType firstType, GgufTensorType secondType) { - if (firstType == GgufTensorType.Q4_0) { - return secondType == GgufTensorType.Q4_0; - } - return firstType == GgufTensorType.Q4_K && secondType == GgufTensorType.Q4_K; - } - - @Override - public boolean isDualEligible( - GgufTensorType firstType, - int firstRows, - GgufTensorType secondType, - int secondRows, - int batchSize, - int cols) { - return supportsDual(firstType, secondType) - && cols % firstType.blockSize() == 0 - && deviceAddressable(firstType, firstRows, cols) - && deviceAddressable(secondType, secondRows, cols) - && eligibleCombined(batchSize, cols, firstRows, secondRows); - } - - @Override - public synchronized void multiplyDual( - float[] firstOutput, - MemorySegment firstWeights, - GgufTensorType firstType, - int firstRows, - float[] secondOutput, - MemorySegment secondWeights, - GgufTensorType secondType, - int secondRows, - float[] input, - int batchSize, - int cols) { - requireOpen(); - if (!isDualEligible(firstType, firstRows, secondType, secondRows, batchSize, cols)) { - throw new UnsupportedOperationException( - "dual projection is not eligible for the Tornado backend"); - } - validateProjectionStorage(firstOutput, firstWeights, input, batchSize, firstRows, cols); - validateProjectionStorage(secondOutput, secondWeights, input, batchSize, secondRows, cols); - int planBatchSize = executionBatchSizeFor(batchSize); - long started = System.nanoTime(); - if (isKQuant(firstType)) { - kQuantPlan( - new MemorySegment[] {firstWeights, secondWeights}, - new GgufTensorType[] {firstType, secondType}, - new int[] {firstRows, secondRows}, - planBatchSize, - cols) - .execute(input, new float[][] {firstOutput, secondOutput}, batchSize); - } else { - DualProjectionKey key = - new DualProjectionKey( - firstWeights.address(), - firstWeights.byteSize(), - firstRows, - secondWeights.address(), - secondWeights.byteSize(), - secondRows, - planBatchSize, - cols); - DualProjectionPlan plan = - dualPlans.computeIfAbsent( - key, - ignored -> - new DualProjectionPlan( - nextPlanName(), - firstWeights, - firstRows, - secondWeights, - secondRows, - planBatchSize, - cols)); - plan.execute(input, firstOutput, secondOutput, batchSize); - } - recordCall(started, firstType, secondType); - } - - @Override - public boolean supportsTriple( - GgufTensorType firstType, GgufTensorType secondType, GgufTensorType thirdType) { - if (firstType == GgufTensorType.Q4_0) { - return secondType == GgufTensorType.Q4_0 && thirdType == GgufTensorType.Q4_0; - } - if (firstType != GgufTensorType.Q4_K || secondType != GgufTensorType.Q4_K) { - return false; - } - // Q4_K_M promotes the value projection to Q6_K while query and key stay Q4_K. - return thirdType == GgufTensorType.Q4_K || thirdType == GgufTensorType.Q6_K; - } - - @Override - public boolean isTripleEligible( - GgufTensorType firstType, - int firstRows, - GgufTensorType secondType, - int secondRows, - GgufTensorType thirdType, - int thirdRows, - int batchSize, - int cols) { - return supportsTriple(firstType, secondType, thirdType) - && cols % firstType.blockSize() == 0 - && cols % thirdType.blockSize() == 0 - && deviceAddressable(firstType, firstRows, cols) - && deviceAddressable(secondType, secondRows, cols) - && deviceAddressable(thirdType, thirdRows, cols) - && eligibleCombined(batchSize, cols, firstRows, secondRows, thirdRows); - } - - @Override - public synchronized void multiplyTriple( - float[] firstOutput, - MemorySegment firstWeights, - GgufTensorType firstType, - int firstRows, - float[] secondOutput, - MemorySegment secondWeights, - GgufTensorType secondType, - int secondRows, - float[] thirdOutput, - MemorySegment thirdWeights, - GgufTensorType thirdType, - int thirdRows, - float[] input, - int batchSize, - int cols) { - requireOpen(); - if (!isTripleEligible( - firstType, firstRows, secondType, secondRows, thirdType, thirdRows, batchSize, cols)) { - throw new UnsupportedOperationException( - "triple projection is not eligible for the Tornado backend"); - } - validateProjectionStorage(firstOutput, firstWeights, input, batchSize, firstRows, cols); - validateProjectionStorage(secondOutput, secondWeights, input, batchSize, secondRows, cols); - validateProjectionStorage(thirdOutput, thirdWeights, input, batchSize, thirdRows, cols); - int planBatchSize = executionBatchSizeFor(batchSize); - long started = System.nanoTime(); - if (isKQuant(firstType)) { - kQuantPlan( - new MemorySegment[] {firstWeights, secondWeights, thirdWeights}, - new GgufTensorType[] {firstType, secondType, thirdType}, - new int[] {firstRows, secondRows, thirdRows}, - planBatchSize, - cols) - .execute(input, new float[][] {firstOutput, secondOutput, thirdOutput}, batchSize); - } else { - TripleProjectionKey key = - new TripleProjectionKey( - firstWeights.address(), - firstWeights.byteSize(), - firstRows, - secondWeights.address(), - secondWeights.byteSize(), - secondRows, - thirdWeights.address(), - thirdWeights.byteSize(), - thirdRows, - planBatchSize, - cols); - TripleProjectionPlan plan = - triplePlans.computeIfAbsent( - key, - ignored -> - new TripleProjectionPlan( - nextPlanName(), - firstWeights, - firstRows, - secondWeights, - secondRows, - thirdWeights, - thirdRows, - planBatchSize, - cols)); - plan.execute(input, firstOutput, secondOutput, thirdOutput, batchSize); - } - recordCall(started, firstType, secondType, thirdType); - } - - @Override - public synchronized void multiply( - float[] output, - float[] input, - MemorySegment weights, - GgufTensorType type, - int batchSize, - int rows, - int cols) { - requireOpen(); - if (!isEligible(type, batchSize, rows, cols)) { - throw new UnsupportedOperationException("projection is not eligible for the Tornado backend"); - } - validateProjectionStorage(output, weights, input, batchSize, rows, cols); - int planBatchSize = executionBatchSizeFor(batchSize); - long started = System.nanoTime(); - if (isKQuant(type)) { - kQuantPlan( - new MemorySegment[] {weights}, - new GgufTensorType[] {type}, - new int[] {rows}, - planBatchSize, - cols) - .execute(input, new float[][] {output}, batchSize); - } else { - ProjectionKey key = - new ProjectionKey(weights.address(), weights.byteSize(), planBatchSize, rows, cols); - ProjectionPlan plan = - plans.computeIfAbsent( - key, - ignored -> new ProjectionPlan(nextPlanName(), weights, planBatchSize, rows, cols)); - plan.execute(input, output, batchSize); - } - recordCall(started, type); - } - - /** Number of distinct tensor/shape execution plans compiled or awaiting first compilation. */ - public synchronized int projectionPlanCount() { - return plans.size() + dualPlans.size() + triplePlans.size() + kQuantPlans.size(); - } - - /** - * Number of model projections routed through this provider, by GGUF weight format. - * - *

A grouped dispatch counts once per matrix, so a Q4_K/Q4_K/Q6_K attention group adds two to - * {@code Q4_K} and one to {@code Q6_K}. Formats absent from the map were never accelerated, which - * is how a run shows that a mixed-format model actually took the device path for each of its - * formats rather than silently falling back for one of them. - */ - public synchronized Map routedProjectionsByFormat() { - Map byFormat = new LinkedHashMap<>(); - routedProjections.forEach((type, count) -> byFormat.put(type.name(), count)); - return Collections.unmodifiableMap(byFormat); - } - - /** Number of model projection calls routed through this provider. */ - public synchronized long calls() { - return calls; - } - - /** Wall-clock time spent preparing, compiling, executing, and copying accelerated calls. */ - public synchronized double totalMillis() { - return totalNanos / 1_000_000.0; - } - - @Override - public synchronized void close() { - if (closed) { - return; - } - RuntimeException failure = null; - for (ProjectionPlan plan : plans.values()) { - try { - plan.close(); - } catch (RuntimeException exception) { - if (failure == null) { - failure = exception; - } else { - failure.addSuppressed(exception); - } - } - } - for (DualProjectionPlan plan : dualPlans.values()) { - try { - plan.close(); - } catch (RuntimeException exception) { - if (failure == null) { - failure = exception; - } else { - failure.addSuppressed(exception); - } - } - } - for (TripleProjectionPlan plan : triplePlans.values()) { - try { - plan.close(); - } catch (RuntimeException exception) { - if (failure == null) { - failure = exception; - } else { - failure.addSuppressed(exception); - } - } - } - for (KQuantProjectionPlan plan : kQuantPlans.values()) { - try { - plan.close(); - } catch (RuntimeException exception) { - if (failure == null) { - failure = exception; - } else { - failure.addSuppressed(exception); - } - } - } - plans.clear(); - dualPlans.clear(); - triplePlans.clear(); - kQuantPlans.clear(); - closed = true; - if (failure != null) { - throw failure; - } - } - - private void requireOpen() { - if (closed) { - throw new IllegalStateException("Tornado projection kernel is closed"); - } - } - - private boolean eligibleCombined(int batchSize, int cols, int... rows) { - if (!eligibleBatchSize(batchSize) || batchSize > executionBatchSize || cols <= 0) { - return false; - } - long combinedRows = 0; - for (int rowCount : rows) { - if (rowCount <= 0) { - return false; - } - combinedRows += rowCount; - } - return combinedRows * cols >= MINIMUM_MATRIX_VALUES; - } - - private boolean eligibleBatchSize(int batchSize) { - return batchSize >= MINIMUM_BATCH || (accelerateDecode && batchSize == 1); - } - - private static void validateProjectionStorage( - float[] output, MemorySegment weights, float[] input, int batchSize, int rows, int cols) { - Objects.requireNonNull(output, "output"); - Objects.requireNonNull(input, "input"); - Objects.requireNonNull(weights, "weights"); - int expectedInput = Math.multiplyExact(batchSize, cols); - int expectedOutput = Math.multiplyExact(batchSize, rows); - if (input.length < expectedInput || output.length < expectedOutput) { - throw new IllegalArgumentException("input or output storage does not match projection shape"); - } - } - - private String nextPlanName() { - return "q4-model-" + planSequence++; - } - - private void recordCall(long started, GgufTensorType... types) { - totalNanos += System.nanoTime() - started; - calls++; - for (GgufTensorType type : types) { - routedProjections.merge(type, 1L, Long::sum); - } - } - - private static boolean isKQuant(GgufTensorType type) { - return type == GgufTensorType.Q4_K || type == GgufTensorType.Q6_K; - } - - private static boolean deviceAddressable(GgufTensorType type, int rows, int cols) { - if (rows <= 0 || cols <= 0 || cols % type.blockSize() != 0) { - return false; - } - return (long) rows * (cols / type.blockSize()) * type.typeSize() <= MAX_DEVICE_TENSOR_BYTES; - } - - private KQuantProjectionPlan kQuantPlan( - MemorySegment[] weights, GgufTensorType[] types, int[] rows, int planBatchSize, int cols) { - List addresses = new ArrayList<>(weights.length); - List byteSizes = new ArrayList<>(weights.length); - List rowCounts = new ArrayList<>(rows.length); - for (MemorySegment weight : weights) { - addresses.add(weight.address()); - byteSizes.add(weight.byteSize()); - } - for (int rowCount : rows) { - rowCounts.add(rowCount); - } - KQuantProjectionKey key = - new KQuantProjectionKey( - List.copyOf(addresses), - List.copyOf(byteSizes), - List.copyOf(rowCounts), - List.of(types), - planBatchSize, - cols); - return kQuantPlans.computeIfAbsent( - key, - ignored -> - new KQuantProjectionPlan(nextPlanName(), weights, types, rows, planBatchSize, cols)); - } - - private record ProjectionKey(long address, long weightBytes, int batchSize, int rows, int cols) {} - - private record DualProjectionKey( - long firstAddress, - long firstWeightBytes, - int firstRows, - long secondAddress, - long secondWeightBytes, - int secondRows, - int batchSize, - int cols) {} - - private record TripleProjectionKey( - long firstAddress, - long firstWeightBytes, - int firstRows, - long secondAddress, - long secondWeightBytes, - int secondRows, - long thirdAddress, - long thirdWeightBytes, - int thirdRows, - int batchSize, - int cols) {} - - private static final class ProjectionPlan implements AutoCloseable { - private final int batchSize; - private final int rows; - private final int cols; - private final float[] paddedInput; - private final byte[] preparedActivations; - private final float[] preparedScales; - private final ByteArray deviceActivations; - private final FloatArray deviceScales; - private final FloatArray deviceOutput; - private final TornadoExecutionPlan plan; - - private ProjectionPlan(String name, MemorySegment weights, int batchSize, int rows, int cols) { - this.batchSize = batchSize; - this.rows = rows; - this.cols = cols; - int activationEntries = Math.multiplyExact(batchSize, cols); - this.paddedInput = new float[activationEntries]; - this.preparedActivations = new byte[activationEntries]; - this.preparedScales = new float[activationEntries / 32]; - ByteArray deviceWeights = ByteArray.fromSegment(weights); - this.deviceActivations = new ByteArray(activationEntries); - this.deviceScales = new FloatArray(preparedScales.length); - this.deviceOutput = new FloatArray(Math.multiplyExact(batchSize, rows)); - Q4ProjectionKernel.validate( - deviceWeights, deviceActivations, deviceScales, deviceOutput, batchSize, rows, cols); - TaskGraph graph = - new TaskGraph(name) - .transferToDevice(DataTransferMode.FIRST_EXECUTION, deviceWeights) - .transferToDevice(DataTransferMode.EVERY_EXECUTION, deviceActivations, deviceScales) - .task( - "multiply", - Q4ProjectionKernel::multiply, - deviceWeights, - deviceActivations, - deviceScales, - deviceOutput, - batchSize, - rows, - cols) - .transferToHost(DataTransferMode.EVERY_EXECUTION, deviceOutput); - this.plan = new TornadoExecutionPlan(graph.snapshot()); - } - - private void execute(float[] input, float[] output, int actualBatchSize) { - float[] executionInput = - prepareExecutionInput(input, paddedInput, actualBatchSize, batchSize, cols); - Q4ProjectionKernel.quantize( - executionInput, preparedActivations, preparedScales, batchSize, cols); - deviceActivations.getSegment().copyFrom(MemorySegment.ofArray(preparedActivations)); - deviceScales.getSegment().copyFrom(MemorySegment.ofArray(preparedScales)); - plan.execute(); - copyOutput(deviceOutput, output, actualBatchSize, rows); - } - - @Override - public void close() { - try { - plan.close(); - } catch (Exception exception) { - throw new IllegalStateException("could not close Tornado projection plan", exception); - } - } - } - - private static final class DualProjectionPlan implements AutoCloseable { - private final int batchSize; - private final int firstRows; - private final int secondRows; - private final int cols; - private final float[] paddedInput; - private final byte[] preparedActivations; - private final float[] preparedScales; - private final ByteArray deviceActivations; - private final FloatArray deviceScales; - private final FloatArray firstDeviceOutput; - private final FloatArray secondDeviceOutput; - private final TornadoExecutionPlan plan; - - private DualProjectionPlan( - String name, - MemorySegment firstWeights, - int firstRows, - MemorySegment secondWeights, - int secondRows, - int batchSize, - int cols) { - this.batchSize = batchSize; - this.firstRows = firstRows; - this.secondRows = secondRows; - this.cols = cols; - int activationEntries = Math.multiplyExact(batchSize, cols); - this.paddedInput = new float[activationEntries]; - this.preparedActivations = new byte[activationEntries]; - this.preparedScales = new float[activationEntries / 32]; - ByteArray firstDeviceWeights = ByteArray.fromSegment(firstWeights); - ByteArray secondDeviceWeights = ByteArray.fromSegment(secondWeights); - this.deviceActivations = new ByteArray(activationEntries); - this.deviceScales = new FloatArray(preparedScales.length); - this.firstDeviceOutput = new FloatArray(Math.multiplyExact(batchSize, firstRows)); - this.secondDeviceOutput = new FloatArray(Math.multiplyExact(batchSize, secondRows)); - Q4ProjectionKernel.validate( - firstDeviceWeights, - deviceActivations, - deviceScales, - firstDeviceOutput, - batchSize, - firstRows, - cols); - Q4ProjectionKernel.validate( - secondDeviceWeights, - deviceActivations, - deviceScales, - secondDeviceOutput, - batchSize, - secondRows, - cols); - TaskGraph graph = - new TaskGraph(name) - .transferToDevice( - DataTransferMode.FIRST_EXECUTION, firstDeviceWeights, secondDeviceWeights) - .transferToDevice(DataTransferMode.EVERY_EXECUTION, deviceActivations, deviceScales) - .task( - "multiply-dual", - Q4ProjectionKernel::multiplyDual, - firstDeviceWeights, - firstRows, - secondDeviceWeights, - secondRows, - deviceActivations, - deviceScales, - firstDeviceOutput, - secondDeviceOutput, - batchSize, - cols) - .transferToHost( - DataTransferMode.EVERY_EXECUTION, firstDeviceOutput, secondDeviceOutput); - this.plan = new TornadoExecutionPlan(graph.snapshot()); - } - - private void execute( - float[] input, float[] firstOutput, float[] secondOutput, int actualBatchSize) { - float[] executionInput = - prepareExecutionInput(input, paddedInput, actualBatchSize, batchSize, cols); - prepareAndStage( - executionInput, - preparedActivations, - preparedScales, - deviceActivations, - deviceScales, - batchSize, - cols); - plan.execute(); - copyOutput(firstDeviceOutput, firstOutput, actualBatchSize, firstRows); - copyOutput(secondDeviceOutput, secondOutput, actualBatchSize, secondRows); - } - - @Override - public void close() { - closePlan(plan); - } - } - - private static final class TripleProjectionPlan implements AutoCloseable { - private final int batchSize; - private final int firstRows; - private final int secondRows; - private final int thirdRows; - private final int cols; - private final float[] paddedInput; - private final byte[] preparedActivations; - private final float[] preparedScales; - private final ByteArray deviceActivations; - private final FloatArray deviceScales; - private final FloatArray firstDeviceOutput; - private final FloatArray secondDeviceOutput; - private final FloatArray thirdDeviceOutput; - private final TornadoExecutionPlan plan; - - private TripleProjectionPlan( - String name, - MemorySegment firstWeights, - int firstRows, - MemorySegment secondWeights, - int secondRows, - MemorySegment thirdWeights, - int thirdRows, - int batchSize, - int cols) { - this.batchSize = batchSize; - this.firstRows = firstRows; - this.secondRows = secondRows; - this.thirdRows = thirdRows; - this.cols = cols; - int activationEntries = Math.multiplyExact(batchSize, cols); - this.paddedInput = new float[activationEntries]; - this.preparedActivations = new byte[activationEntries]; - this.preparedScales = new float[activationEntries / 32]; - ByteArray firstDeviceWeights = ByteArray.fromSegment(firstWeights); - ByteArray secondDeviceWeights = ByteArray.fromSegment(secondWeights); - ByteArray thirdDeviceWeights = ByteArray.fromSegment(thirdWeights); - this.deviceActivations = new ByteArray(activationEntries); - this.deviceScales = new FloatArray(preparedScales.length); - this.firstDeviceOutput = new FloatArray(Math.multiplyExact(batchSize, firstRows)); - this.secondDeviceOutput = new FloatArray(Math.multiplyExact(batchSize, secondRows)); - this.thirdDeviceOutput = new FloatArray(Math.multiplyExact(batchSize, thirdRows)); - Q4ProjectionKernel.validate( - firstDeviceWeights, - deviceActivations, - deviceScales, - firstDeviceOutput, - batchSize, - firstRows, - cols); - Q4ProjectionKernel.validate( - secondDeviceWeights, - deviceActivations, - deviceScales, - secondDeviceOutput, - batchSize, - secondRows, - cols); - Q4ProjectionKernel.validate( - thirdDeviceWeights, - deviceActivations, - deviceScales, - thirdDeviceOutput, - batchSize, - thirdRows, - cols); - TaskGraph graph = - new TaskGraph(name) - .transferToDevice( - DataTransferMode.FIRST_EXECUTION, - firstDeviceWeights, - secondDeviceWeights, - thirdDeviceWeights) - .transferToDevice(DataTransferMode.EVERY_EXECUTION, deviceActivations, deviceScales) - .task( - "multiply-triple", - Q4ProjectionKernel::multiplyTriple, - firstDeviceWeights, - firstRows, - secondDeviceWeights, - secondRows, - thirdDeviceWeights, - thirdRows, - deviceActivations, - deviceScales, - firstDeviceOutput, - secondDeviceOutput, - thirdDeviceOutput, - batchSize, - cols) - .transferToHost( - DataTransferMode.EVERY_EXECUTION, - firstDeviceOutput, - secondDeviceOutput, - thirdDeviceOutput); - this.plan = new TornadoExecutionPlan(graph.snapshot()); - } - - private void execute( - float[] input, - float[] firstOutput, - float[] secondOutput, - float[] thirdOutput, - int actualBatchSize) { - float[] executionInput = - prepareExecutionInput(input, paddedInput, actualBatchSize, batchSize, cols); - prepareAndStage( - executionInput, - preparedActivations, - preparedScales, - deviceActivations, - deviceScales, - batchSize, - cols); - plan.execute(); - copyOutput(firstDeviceOutput, firstOutput, actualBatchSize, firstRows); - copyOutput(secondDeviceOutput, secondOutput, actualBatchSize, secondRows); - copyOutput(thirdDeviceOutput, thirdOutput, actualBatchSize, thirdRows); - } - - @Override - public void close() { - closePlan(plan); - } - } - - private record KQuantProjectionKey( - List addresses, - List byteSizes, - List rows, - List types, - int batchSize, - int cols) {} - - /** - * One compiled K-quant dispatch: up to three weight tensors sharing a single Q8_K activation - * preparation. - * - *

The shape is fixed at construction so the task graph and its compiled code are reused across - * calls, exactly like the Q4_0 plans. Only shapes {@link #supportsDual} and {@link - * #supportsTriple} admit can reach here, so the task selection below is total. - */ - private static final class KQuantProjectionPlan implements AutoCloseable { - private final int batchSize; - private final int cols; - private final int[] rows; - private final float[] paddedInput; - private final byte[] preparedActivations; - private final float[] preparedScales; - private final int[] preparedSums; - private final ByteArray deviceActivations; - private final FloatArray deviceScales; - private final IntArray deviceSums; - private final FloatArray[] deviceOutputs; - private final boolean stagesSums; - private final TornadoExecutionPlan plan; - - private KQuantProjectionPlan( - String name, - MemorySegment[] weights, - GgufTensorType[] types, - int[] rows, - int batchSize, - int cols) { - this.batchSize = batchSize; - this.cols = cols; - this.rows = rows.clone(); - int activationEntries = Math.multiplyExact(batchSize, cols); - this.paddedInput = new float[activationEntries]; - this.preparedActivations = new byte[activationEntries]; - this.preparedScales = - new float[activationEntries / KQuantProjectionKernel.SUPER_BLOCK_VALUES]; - this.preparedSums = new int[activationEntries / KQuantProjectionKernel.SUM_BLOCK_VALUES]; - this.deviceActivations = new ByteArray(activationEntries); - this.deviceScales = new FloatArray(preparedScales.length); - this.deviceSums = new IntArray(preparedSums.length); - ByteArray[] deviceWeights = new ByteArray[weights.length]; - this.deviceOutputs = new FloatArray[weights.length]; - for (int index = 0; index < weights.length; index++) { - deviceWeights[index] = ByteArray.fromSegment(weights[index]); - deviceOutputs[index] = new FloatArray(Math.multiplyExact(batchSize, rows[index])); - if (types[index] == GgufTensorType.Q6_K) { - KQuantProjectionKernel.validateQ6K( - deviceWeights[index], - deviceActivations, - deviceScales, - deviceOutputs[index], - batchSize, - rows[index], - cols); - } else { - KQuantProjectionKernel.validateQ4K( - deviceWeights[index], - deviceActivations, - deviceScales, - deviceSums, - deviceOutputs[index], - batchSize, - rows[index], - cols); - } - } - // A Q6_K-only dispatch reads no Q8_K block sums, so the buffer must not be staged for a - // task that never takes it as a parameter. - boolean usesSums = false; - for (GgufTensorType type : types) { - usesSums |= type == GgufTensorType.Q4_K; - } - this.stagesSums = usesSums; - TaskGraph graph = - new TaskGraph(name) - .transferToDevice(DataTransferMode.FIRST_EXECUTION, (Object[]) deviceWeights); - graph = - usesSums - ? graph.transferToDevice( - DataTransferMode.EVERY_EXECUTION, deviceActivations, deviceScales, deviceSums) - : graph.transferToDevice( - DataTransferMode.EVERY_EXECUTION, deviceActivations, deviceScales); - this.plan = - new TornadoExecutionPlan( - addTask(graph, deviceWeights, types, rows) - .transferToHost(DataTransferMode.EVERY_EXECUTION, (Object[]) deviceOutputs) - .snapshot()); - } - - private TaskGraph addTask( - TaskGraph graph, ByteArray[] deviceWeights, GgufTensorType[] types, int[] rows) { - if (deviceWeights.length == 1) { - if (types[0] == GgufTensorType.Q6_K) { - return graph.task( - "multiply-q6k", - KQuantProjectionKernel::multiplyQ6K, - deviceWeights[0], - deviceActivations, - deviceScales, - deviceOutputs[0], - batchSize, - rows[0], - cols); - } - return graph.task( - "multiply-q4k", - KQuantProjectionKernel::multiplyQ4K, - deviceWeights[0], - deviceActivations, - deviceScales, - deviceSums, - deviceOutputs[0], - batchSize, - rows[0], - cols); - } - if (deviceWeights.length == 2) { - return graph.task( - "multiply-q4k-dual", - KQuantProjectionKernel::multiplyQ4KDual, - deviceWeights[0], - rows[0], - deviceWeights[1], - rows[1], - deviceActivations, - deviceScales, - deviceSums, - deviceOutputs[0], - deviceOutputs[1], - batchSize, - cols); - } - if (types[2] == GgufTensorType.Q6_K) { - return graph.task( - "multiply-q4k-q4k-q6k", - KQuantProjectionKernel::multiplyMixedTriple, - deviceWeights[0], - rows[0], - deviceWeights[1], - rows[1], - deviceWeights[2], - rows[2], - deviceActivations, - deviceScales, - deviceSums, - deviceOutputs[0], - deviceOutputs[1], - deviceOutputs[2], - batchSize, - cols); - } - return graph.task( - "multiply-q4k-triple", - KQuantProjectionKernel::multiplyQ4KTriple, - deviceWeights[0], - rows[0], - deviceWeights[1], - rows[1], - deviceWeights[2], - rows[2], - deviceActivations, - deviceScales, - deviceSums, - deviceOutputs[0], - deviceOutputs[1], - deviceOutputs[2], - batchSize, - cols); - } - - private void execute(float[] input, float[][] outputs, int actualBatchSize) { - float[] executionInput = - prepareExecutionInput(input, paddedInput, actualBatchSize, batchSize, cols); - KQuantProjectionKernel.quantize( - executionInput, preparedActivations, preparedScales, preparedSums, batchSize, cols); - deviceActivations.getSegment().copyFrom(MemorySegment.ofArray(preparedActivations)); - deviceScales.getSegment().copyFrom(MemorySegment.ofArray(preparedScales)); - if (stagesSums) { - deviceSums.getSegment().copyFrom(MemorySegment.ofArray(preparedSums)); - } - plan.execute(); - for (int index = 0; index < outputs.length; index++) { - copyOutput(deviceOutputs[index], outputs[index], actualBatchSize, rows[index]); - } - } - - @Override - public void close() { - closePlan(plan); - } - } - - private static void prepareAndStage( - float[] input, - byte[] preparedActivations, - float[] preparedScales, - ByteArray deviceActivations, - FloatArray deviceScales, - int batchSize, - int cols) { - Q4ProjectionKernel.quantize(input, preparedActivations, preparedScales, batchSize, cols); - deviceActivations.getSegment().copyFrom(MemorySegment.ofArray(preparedActivations)); - deviceScales.getSegment().copyFrom(MemorySegment.ofArray(preparedScales)); - } - - static float[] prepareExecutionInput( - float[] input, float[] paddedInput, int actualBatchSize, int executionBatchSize, int cols) { - int actualEntries = Math.multiplyExact(actualBatchSize, cols); - int executionEntries = Math.multiplyExact(executionBatchSize, cols); - if (input.length < actualEntries || paddedInput.length != executionEntries) { - throw new IllegalArgumentException("input storage does not match execution batch shape"); - } - System.arraycopy(input, 0, paddedInput, 0, actualEntries); - Arrays.fill(paddedInput, actualEntries, executionEntries, 0.0f); - return paddedInput; - } - - private static void copyOutput(FloatArray deviceOutput, float[] output, int batchSize, int rows) { - long byteSize = Math.multiplyExact((long) batchSize * rows, Float.BYTES); - MemorySegment.ofArray(output) - .asSlice(0, byteSize) - .copyFrom(deviceOutput.getSegment().asSlice(0, byteSize)); - } - - private static void closePlan(TornadoExecutionPlan plan) { - try { - plan.close(); - } catch (Exception exception) { - throw new IllegalStateException("could not close Tornado projection plan", exception); - } - } -} diff --git a/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/TornadoPureJavaBackendProvider.java b/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/TornadoPureJavaBackendProvider.java deleted file mode 100644 index 4937bedbb..000000000 --- a/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/TornadoPureJavaBackendProvider.java +++ /dev/null @@ -1,72 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.backend.tornado; - -import com.integrallis.models.api.BackendConfiguration; -import com.integrallis.models.backend.purejava.PureJavaBackend; -import com.integrallis.models.backend.purejava.spi.PureJavaBackendProvider; -import java.io.IOException; -import java.nio.file.Files; -import java.nio.file.Path; -import java.util.List; -import java.util.Optional; - -/** Service-loaded Tornado provider used by Models and ModelJars automatic backend loading. */ -public final class TornadoPureJavaBackendProvider implements PureJavaBackendProvider { - - @Override - public Optional tryLoad( - Path modelPath, BackendConfiguration backendConfiguration) { - Path model = modelPath.toAbsolutePath().normalize(); - if (!Files.isRegularFile(model)) { - return Optional.empty(); - } - List devices; - try { - devices = TornadoRuntimeDevices.discover(); - } catch (LinkageError | RuntimeException failure) { - return Optional.empty(); - } - long modelBytes; - try { - modelBytes = Files.size(model); - } catch (IOException failure) { - return Optional.empty(); - } - TornadoBackendOptions options = TornadoBackendOptions.defaults(); - AcceleratorEligibility.Decision decision = - AcceleratorEligibility.select( - devices, - DeviceMemoryRequest.ofModelFile( - modelLabel(model), modelBytes, options.accelerateDecode())); - if (!decision.eligible()) { - return Optional.empty(); - } - TornadoBackendOptions required = - new TornadoBackendOptions( - options.accelerateDecode(), - options.eagerReadiness(), - true, - options.executionBatchSize()); - TornadoBackendRuntime runtime = TornadoBackend.open(model, backendConfiguration, required); - return Optional.of(runtime.detachBackend()); - } - - private static String modelLabel(Path model) { - Path fileName = model.getFileName(); - return fileName == null ? model.toString() : fileName.toString(); - } -} diff --git a/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/TornadoRuntimeDevices.java b/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/TornadoRuntimeDevices.java deleted file mode 100644 index 90200047c..000000000 --- a/backend-tornado/src/main/java/com/integrallis/models/backend/tornado/TornadoRuntimeDevices.java +++ /dev/null @@ -1,52 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.backend.tornado; - -import java.util.ArrayList; -import java.util.List; -import uk.ac.manchester.tornado.api.TornadoBackend; -import uk.ac.manchester.tornado.api.TornadoRuntime; -import uk.ac.manchester.tornado.api.common.TornadoDevice; -import uk.ac.manchester.tornado.api.runtime.TornadoRuntimeProvider; - -/** Discovers accelerator devices from the active TornadoVM runtime. */ -public final class TornadoRuntimeDevices { - private TornadoRuntimeDevices() {} - - public static List discover() { - try { - TornadoRuntime runtime = TornadoRuntimeProvider.getTornadoRuntime(); - List devices = new ArrayList<>(); - for (int backendIndex = 0; backendIndex < runtime.getNumBackends(); backendIndex++) { - TornadoBackend backend = runtime.getBackend(backendIndex); - String backendType = runtime.getBackendType(backendIndex).name(); - for (int deviceIndex = 0; deviceIndex < backend.getNumDevices(); deviceIndex++) { - TornadoDevice device = backend.getDevice(deviceIndex); - devices.add( - new AcceleratorEligibility.DeviceCapabilities( - device.getDeviceName(), - backendType, - device.getDeviceType().name(), - device.getMaxGlobalMemory(), - device.getMaxAllocMemory())); - } - } - return List.copyOf(devices); - } catch (LinkageError | RuntimeException unavailable) { - return List.of(); - } - } -} diff --git a/backend-tornado/src/main/resources/META-INF/services/com.integrallis.models.backend.purejava.spi.PureJavaBackendProvider b/backend-tornado/src/main/resources/META-INF/services/com.integrallis.models.backend.purejava.spi.PureJavaBackendProvider deleted file mode 100644 index adddedcdf..000000000 --- a/backend-tornado/src/main/resources/META-INF/services/com.integrallis.models.backend.purejava.spi.PureJavaBackendProvider +++ /dev/null @@ -1 +0,0 @@ -com.integrallis.models.backend.tornado.TornadoPureJavaBackendProvider diff --git a/backend-tornado/src/test/java/com/integrallis/models/backend/tornado/AcceleratorEligibilityTest.java b/backend-tornado/src/test/java/com/integrallis/models/backend/tornado/AcceleratorEligibilityTest.java deleted file mode 100644 index 7ad81d3e1..000000000 --- a/backend-tornado/src/test/java/com/integrallis/models/backend/tornado/AcceleratorEligibilityTest.java +++ /dev/null @@ -1,80 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.backend.tornado; - -import static org.assertj.core.api.Assertions.assertThat; - -import java.util.List; -import org.junit.jupiter.api.Test; - -class AcceleratorEligibilityTest { - private static final long MIB = 1024L * 1024L; - private static final long GIB = 1024L * MIB; - - @Test - void selectsAQualifiedPtxGpuWithEnoughMemoryForPrefillAndDecodePlans() { - AcceleratorEligibility.Decision decision = - AcceleratorEligibility.select( - List.of( - device("host CPU", "JAVA", "CPU", 16 * GIB, 4 * GIB), - device("NVIDIA A16", "PTX", "GPU", 2 * GIB, 512 * MIB)), - 410 * MIB, - true); - - assertThat(decision.eligible()).isTrue(); - assertThat(decision.device().name()).isEqualTo("NVIDIA A16"); - assertThat(decision.requiredBytes()).isEqualTo(1_076 * MIB); - } - - @Test - void fallsBackWhenTheQualifiedDeviceCannotHoldTheRetainedPlansWithHeadroom() { - AcceleratorEligibility.Decision decision = - AcceleratorEligibility.select( - List.of(device("NVIDIA A16", "PTX", "GPU", 2 * GIB, 512 * MIB)), 900 * MIB, true); - - assertThat(decision.eligible()).isFalse(); - assertThat(decision.reason()).contains("NVIDIA A16 has 1.50 GiB usable of 2.00 GiB"); - assertThat(decision.reason()).contains("needs 2.01 GiB"); - assertThat(decision.reason()).contains("weights 1.76 GiB (PER_SHAPE_WHOLE_MODEL)"); - } - - @Test - void leavesUnqualifiedVendorsAndCpuDevicesOnTheJavaFallback() { - AcceleratorEligibility.Decision decision = - AcceleratorEligibility.select( - List.of( - device("AMD Radeon", "OPENCL", "GPU", 8 * GIB, 2 * GIB), - device("Intel Xeon", "OPENCL", "CPU", 64 * GIB, 16 * GIB)), - 410 * MIB, - false); - - assertThat(decision.eligible()).isFalse(); - assertThat(decision.reason()).contains("PTX or CUDA backend on a GPU device"); - } - - @Test - void rejectsInvalidModelSizesWithoutInspectingDevices() { - AcceleratorEligibility.Decision decision = AcceleratorEligibility.select(List.of(), 0, false); - - assertThat(decision.eligible()).isFalse(); - assertThat(decision.reason()).contains("model size"); - } - - private static AcceleratorEligibility.DeviceCapabilities device( - String name, String backend, String type, long memory, long allocation) { - return new AcceleratorEligibility.DeviceCapabilities(name, backend, type, memory, allocation); - } -} diff --git a/backend-tornado/src/test/java/com/integrallis/models/backend/tornado/KQuantProjectionKernelTest.java b/backend-tornado/src/test/java/com/integrallis/models/backend/tornado/KQuantProjectionKernelTest.java deleted file mode 100644 index 1af78ed2d..000000000 --- a/backend-tornado/src/test/java/com/integrallis/models/backend/tornado/KQuantProjectionKernelTest.java +++ /dev/null @@ -1,697 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.backend.tornado; - -import static org.assertj.core.api.Assertions.assertThat; -import static org.assertj.core.api.Assertions.assertThatThrownBy; -import static org.assertj.core.api.Assertions.within; - -import com.integrallis.vectors.core.VectorUtil; -import java.lang.foreign.MemorySegment; -import java.util.Random; -import org.junit.jupiter.api.Test; -import uk.ac.manchester.tornado.api.types.arrays.ByteArray; -import uk.ac.manchester.tornado.api.types.arrays.FloatArray; -import uk.ac.manchester.tornado.api.types.arrays.IntArray; - -/** - * Off-device parity tests for the K-quant TornadoVM kernels. - * - *

These run the kernels as ordinary sequential Java — {@code @Parallel} is an annotation, so a - * direct call from a unit test executes the same arithmetic the PTX backend compiles — and score - * them against the production vectors-core CPU kernels the pure-Java backend actually uses. - * - *

Two numeric contracts are asserted, deliberately separately: - * - *

    - *
  1. Exact. On super-blocks whose scales are powers of two and whose quantized sums stay - * well inside the exactly representable float range, the kernel must equal the CPU kernel - * bit-for-bit. This is the strong statement: it proves the GGUF super-block decode — the - * six-bit scale/min packing, the nibble group order, the Q6_K two-bit high pairs, and the - * Q8_K per-16 sums — is identical, because every integer reduction is exact and any decode - * error would move the result by at least one quantization step. - *
  2. Bounded. On pseudo-random super-blocks the kernel cannot be exact: the CPU kernels - * fuse the per-block scale application with {@code Math.fma} and this kernel uses a plain - * multiply and add, which is the operation set the existing Q4_0 kernel is known to compile - * to PTX with. Q4_K applies two such steps per super-block and Q6_K one, so the guaranteed - * contract is at most two float roundings per super-block: {@code |kernel - cpu| <= 2 - * * (cols / 256) * ulp(max |cpu|)}. That is what {@link #assertWithinFusedMultiplyAddBudget} - * asserts. - *

    Measured here on 2026-09-18 (Apple M-series, vectors-core 0.1.22 Panama provider, - * 256-bit species, {@code fastVectorFMA=true}), over 40 pseudo-random matrices per width at - * batch 4 by 8 rows: Q4_K worst absolute difference 7.6e-6 at cols=256, 6.1e-5 at cols=1024, - * 1.2e-4 at cols=4096; Q6_K 0.0, 9.2e-5, 1.8e-4 for the same widths, against reference - * magnitudes of 3.3e2, 5.3e2 and 1.4e3. Every one of those is inside the asserted budget with - * at least an order of magnitude to spare, and all of them scale with the super-block count - * exactly as a per-super-block rounding predicts. - *

- * - *

What these tests cannot establish: that TornadoVM's PTX backend lowers this bytecode to - * the same arithmetic. A device run can differ by contracting {@code a * b + c} into a hardware - * fused multiply-add, which would move results toward the CPU kernel rather than away from it, but - * that is an argument, not a measurement. The device-side statement needs a GPU host. - */ -class KQuantProjectionKernelTest { - private static final int SUPER_BLOCK = KQuantProjectionKernel.SUPER_BLOCK_VALUES; - private static final int SUM_BLOCK = KQuantProjectionKernel.SUM_BLOCK_VALUES; - - // --- Q4_K --- - - @Test - void matchesProductionQ4KProjectionAcrossBatches() { - int batchSize = 3; - int rows = 9; - int cols = 512; - byte[] weights = randomQ4KMatrix(rows, cols, 31L); - float[] input = randomFloats(batchSize * cols, 37L); - - float[] actual = runQ4K(weights, input, batchSize, rows, cols); - - assertWithinFusedMultiplyAddBudget( - actual, referenceQ4K(weights, input, batchSize, rows, cols), cols); - } - - @Test - void matchesProductionQ4KProjectionExactlyOnExactlyRepresentableBlocks() { - int batchSize = 2; - int rows = 5; - int cols = 512; - byte[] weights = exactQ4KMatrix(rows, cols, 71L); - float[] input = exactFloats(batchSize, cols, 73L); - - float[] actual = runQ4K(weights, input, batchSize, rows, cols); - - assertThat(actual).isEqualTo(referenceQ4K(weights, input, batchSize, rows, cols)); - } - - @Test - void matchesTwoProductionQ4KProjectionsWithOnePreparedActivation() { - int batchSize = 2; - int firstRows = 7; - int secondRows = 5; - int cols = 256; - byte[] firstWeights = exactQ4KMatrix(firstRows, cols, 41L); - byte[] secondWeights = exactQ4KMatrix(secondRows, cols, 43L); - float[] input = exactFloats(batchSize, cols, 47L); - Q8KActivation activation = Q8KActivation.of(input, batchSize, cols); - FloatArray firstOutput = new FloatArray(batchSize * firstRows); - FloatArray secondOutput = new FloatArray(batchSize * secondRows); - - KQuantProjectionKernel.multiplyQ4KDual( - ByteArray.fromArray(firstWeights), - firstRows, - ByteArray.fromArray(secondWeights), - secondRows, - activation.quants(), - activation.scales(), - activation.sums(), - firstOutput, - secondOutput, - batchSize, - cols); - - assertThat(firstOutput.toHeapArray()) - .isEqualTo(referenceQ4K(firstWeights, input, batchSize, firstRows, cols)); - assertThat(secondOutput.toHeapArray()) - .isEqualTo(referenceQ4K(secondWeights, input, batchSize, secondRows, cols)); - } - - @Test - void matchesThreeProductionQ4KProjectionsWithOnePreparedActivation() { - int batchSize = 2; - int firstRows = 4; - int secondRows = 3; - int thirdRows = 2; - int cols = 256; - byte[] firstWeights = randomQ4KMatrix(firstRows, cols, 53L); - byte[] secondWeights = randomQ4KMatrix(secondRows, cols, 59L); - byte[] thirdWeights = randomQ4KMatrix(thirdRows, cols, 61L); - float[] input = randomFloats(batchSize * cols, 67L); - Q8KActivation activation = Q8KActivation.of(input, batchSize, cols); - FloatArray firstOutput = new FloatArray(batchSize * firstRows); - FloatArray secondOutput = new FloatArray(batchSize * secondRows); - FloatArray thirdOutput = new FloatArray(batchSize * thirdRows); - - KQuantProjectionKernel.multiplyQ4KTriple( - ByteArray.fromArray(firstWeights), - firstRows, - ByteArray.fromArray(secondWeights), - secondRows, - ByteArray.fromArray(thirdWeights), - thirdRows, - activation.quants(), - activation.scales(), - activation.sums(), - firstOutput, - secondOutput, - thirdOutput, - batchSize, - cols); - - assertWithinFusedMultiplyAddBudget( - firstOutput.toHeapArray(), - referenceQ4K(firstWeights, input, batchSize, firstRows, cols), - cols); - assertWithinFusedMultiplyAddBudget( - secondOutput.toHeapArray(), - referenceQ4K(secondWeights, input, batchSize, secondRows, cols), - cols); - assertWithinFusedMultiplyAddBudget( - thirdOutput.toHeapArray(), - referenceQ4K(thirdWeights, input, batchSize, thirdRows, cols), - cols); - } - - // --- Q6_K --- - - @Test - void matchesProductionQ6KProjectionAcrossBatches() { - int batchSize = 3; - int rows = 9; - int cols = 512; - byte[] weights = randomQ6KMatrix(rows, cols, 83L); - float[] input = randomFloats(batchSize * cols, 89L); - - float[] actual = runQ6K(weights, input, batchSize, rows, cols); - - assertWithinFusedMultiplyAddBudget( - actual, referenceQ6K(weights, input, batchSize, rows, cols), cols); - } - - @Test - void matchesProductionQ6KProjectionExactlyOnExactlyRepresentableBlocks() { - int batchSize = 2; - int rows = 5; - int cols = 512; - byte[] weights = exactQ6KMatrix(rows, cols, 97L); - float[] input = exactFloats(batchSize, cols, 101L); - - float[] actual = runQ6K(weights, input, batchSize, rows, cols); - - assertThat(actual).isEqualTo(referenceQ6K(weights, input, batchSize, rows, cols)); - } - - // --- Mixed Q4_K / Q4_K / Q6_K, the Q4_K_M attention projection group --- - - @Test - void matchesTheProductionMixedAttentionProjectionGroup() { - int batchSize = 2; - int queryRows = 6; - int keyRows = 3; - int valueRows = 3; - int cols = 256; - byte[] queryWeights = exactQ4KMatrix(queryRows, cols, 103L); - byte[] keyWeights = exactQ4KMatrix(keyRows, cols, 107L); - byte[] valueWeights = exactQ6KMatrix(valueRows, cols, 109L); - float[] input = exactFloats(batchSize, cols, 113L); - Q8KActivation activation = Q8KActivation.of(input, batchSize, cols); - FloatArray queryOutput = new FloatArray(batchSize * queryRows); - FloatArray keyOutput = new FloatArray(batchSize * keyRows); - FloatArray valueOutput = new FloatArray(batchSize * valueRows); - - KQuantProjectionKernel.multiplyMixedTriple( - ByteArray.fromArray(queryWeights), - queryRows, - ByteArray.fromArray(keyWeights), - keyRows, - ByteArray.fromArray(valueWeights), - valueRows, - activation.quants(), - activation.scales(), - activation.sums(), - queryOutput, - keyOutput, - valueOutput, - batchSize, - cols); - - assertThat(queryOutput.toHeapArray()) - .isEqualTo(referenceQ4K(queryWeights, input, batchSize, queryRows, cols)); - assertThat(keyOutput.toHeapArray()) - .isEqualTo(referenceQ4K(keyWeights, input, batchSize, keyRows, cols)); - assertThat(valueOutput.toHeapArray()) - .isEqualTo(referenceQ6K(valueWeights, input, batchSize, valueRows, cols)); - } - - // --- Q8_K activation preparation --- - - @Test - void preparesQ8KActivationsWithBlockSumsThatMatchTheStoredQuants() { - int batchSize = 2; - int cols = 512; - float[] input = randomFloats(batchSize * cols, 127L); - byte[] quants = new byte[batchSize * cols]; - float[] scales = new float[batchSize * cols / SUPER_BLOCK]; - int[] sums = new int[batchSize * cols / SUM_BLOCK]; - - KQuantProjectionKernel.quantize(input, quants, scales, sums, batchSize, cols); - - for (int sumBlock = 0; sumBlock < sums.length; sumBlock++) { - int expected = 0; - for (int index = 0; index < SUM_BLOCK; index++) { - expected += quants[sumBlock * SUM_BLOCK + index]; - } - assertThat(sums[sumBlock]).as("sum block %d", sumBlock).isEqualTo(expected); - assertThat(expected).isBetween(-Short.MAX_VALUE, (int) Short.MAX_VALUE); - } - // Q8_K scales carry the sign of the block extremum, unlike Q8_0's absolute-maximum scale. - for (float scale : scales) { - assertThat(scale).isNotZero().isFinite(); - } - for (int index = 0; index < quants.length; index++) { - assertThat((int) quants[index]).isBetween(-127, 127); - float reconstructed = quants[index] * scales[index / SUPER_BLOCK]; - assertThat(reconstructed) - .as("reconstructed element %d", index) - .isCloseTo(input[index], within(0.02f)); - } - } - - @Test - void preparesZeroActivationBlocksWithoutDividingByZero() { - int cols = 256; - byte[] quants = new byte[cols]; - float[] scales = new float[1]; - int[] sums = new int[cols / SUM_BLOCK]; - - KQuantProjectionKernel.quantize(new float[cols], quants, scales, sums, 1, cols); - - assertThat(scales[0]).isZero(); - assertThat(quants).containsOnly((byte) 0); - assertThat(sums).containsOnly(0); - } - - @Test - void projectsZeroActivationsToZero() { - int rows = 8; - int cols = 256; - byte[] weights = randomQ4KMatrix(rows, cols, 131L); - - float[] actual = runQ4K(weights, new float[cols], 1, rows, cols); - - assertThat(actual).containsOnly(0.0f); - } - - // --- Validation --- - - @Test - void rejectsQ4KActivationSumStorageThatDoesNotMatchTheShape() { - assertThatThrownBy( - () -> - KQuantProjectionKernel.validateQ4K( - new ByteArray(KQuantProjectionKernel.Q4_K_BLOCK_BYTES), - new ByteArray(SUPER_BLOCK), - new FloatArray(1), - new IntArray(3), - new FloatArray(1), - 1, - 1, - SUPER_BLOCK)) - .isInstanceOf(IllegalArgumentException.class) - .hasMessageContaining("activation sums"); - } - - @Test - void rejectsQ4KShapesThatAreNotWholeSuperBlocks() { - assertThatThrownBy( - () -> - KQuantProjectionKernel.validateQ4K( - new ByteArray(KQuantProjectionKernel.Q4_K_BLOCK_BYTES), - new ByteArray(128), - new FloatArray(1), - new IntArray(8), - new FloatArray(1), - 1, - 1, - 128)) - .isInstanceOf(IllegalArgumentException.class) - .hasMessageContaining("multiple of 256"); - } - - @Test - void rejectsQ4KWeightsThatDoNotMatchTheShape() { - assertThatThrownBy( - () -> - KQuantProjectionKernel.validateQ4K( - new ByteArray(KQuantProjectionKernel.Q4_K_BLOCK_BYTES + 1), - new ByteArray(SUPER_BLOCK), - new FloatArray(1), - new IntArray(16), - new FloatArray(1), - 1, - 1, - SUPER_BLOCK)) - .isInstanceOf(IllegalArgumentException.class) - .hasMessageContaining("weights"); - } - - @Test - void rejectsQ4KActivationStorageThatDoesNotMatchTheShape() { - assertThatThrownBy( - () -> - KQuantProjectionKernel.validateQ4K( - new ByteArray(KQuantProjectionKernel.Q4_K_BLOCK_BYTES), - new ByteArray(SUPER_BLOCK - 1), - new FloatArray(1), - new IntArray(16), - new FloatArray(1), - 1, - 1, - SUPER_BLOCK)) - .isInstanceOf(IllegalArgumentException.class) - .hasMessageContaining("activations"); - } - - @Test - void rejectsQ4KActivationScaleStorageThatDoesNotMatchTheShape() { - assertThatThrownBy( - () -> - KQuantProjectionKernel.validateQ4K( - new ByteArray(KQuantProjectionKernel.Q4_K_BLOCK_BYTES), - new ByteArray(SUPER_BLOCK), - new FloatArray(2), - new IntArray(16), - new FloatArray(1), - 1, - 1, - SUPER_BLOCK)) - .isInstanceOf(IllegalArgumentException.class) - .hasMessageContaining("activation scales"); - } - - @Test - void rejectsQ6KOutputStorageThatDoesNotMatchTheShape() { - assertThatThrownBy( - () -> - KQuantProjectionKernel.validateQ6K( - new ByteArray(KQuantProjectionKernel.Q6_K_BLOCK_BYTES), - new ByteArray(SUPER_BLOCK), - new FloatArray(1), - new FloatArray(2), - 1, - 1, - SUPER_BLOCK)) - .isInstanceOf(IllegalArgumentException.class) - .hasMessageContaining("output"); - } - - @Test - void rejectsNonPositiveQ6KShapes() { - assertThatThrownBy( - () -> - KQuantProjectionKernel.validateQ6K( - new ByteArray(KQuantProjectionKernel.Q6_K_BLOCK_BYTES), - new ByteArray(SUPER_BLOCK), - new FloatArray(1), - new FloatArray(1), - 0, - 1, - SUPER_BLOCK)) - .isInstanceOf(IllegalArgumentException.class) - .hasMessageContaining("batchSize"); - assertThatThrownBy( - () -> - KQuantProjectionKernel.validateQ6K( - new ByteArray(KQuantProjectionKernel.Q6_K_BLOCK_BYTES), - new ByteArray(SUPER_BLOCK), - new FloatArray(1), - new FloatArray(1), - 1, - 0, - SUPER_BLOCK)) - .isInstanceOf(IllegalArgumentException.class) - .hasMessageContaining("rows"); - } - - @Test - void acceptsWellFormedQ6KStorage() { - KQuantProjectionKernel.validateQ6K( - new ByteArray(KQuantProjectionKernel.Q6_K_BLOCK_BYTES), - new ByteArray(SUPER_BLOCK), - new FloatArray(1), - new FloatArray(1), - 1, - 1, - SUPER_BLOCK); - } - - // --- Harness --- - - private record Q8KActivation(ByteArray quants, FloatArray scales, IntArray sums) { - static Q8KActivation of(float[] input, int batchSize, int cols) { - byte[] quants = new byte[batchSize * cols]; - float[] scales = new float[batchSize * cols / SUPER_BLOCK]; - int[] sums = new int[batchSize * cols / SUM_BLOCK]; - KQuantProjectionKernel.quantize(input, quants, scales, sums, batchSize, cols); - return new Q8KActivation( - ByteArray.fromArray(quants), FloatArray.fromArray(scales), IntArray.fromArray(sums)); - } - } - - private static float[] runQ4K(byte[] weights, float[] input, int batchSize, int rows, int cols) { - Q8KActivation activation = Q8KActivation.of(input, batchSize, cols); - FloatArray output = new FloatArray(batchSize * rows); - ByteArray deviceWeights = ByteArray.fromArray(weights); - KQuantProjectionKernel.validateQ4K( - deviceWeights, - activation.quants(), - activation.scales(), - activation.sums(), - output, - batchSize, - rows, - cols); - KQuantProjectionKernel.multiplyQ4K( - deviceWeights, - activation.quants(), - activation.scales(), - activation.sums(), - output, - batchSize, - rows, - cols); - return output.toHeapArray(); - } - - private static float[] runQ6K(byte[] weights, float[] input, int batchSize, int rows, int cols) { - Q8KActivation activation = Q8KActivation.of(input, batchSize, cols); - FloatArray output = new FloatArray(batchSize * rows); - ByteArray deviceWeights = ByteArray.fromArray(weights); - KQuantProjectionKernel.validateQ6K( - deviceWeights, activation.quants(), activation.scales(), output, batchSize, rows, cols); - KQuantProjectionKernel.multiplyQ6K( - deviceWeights, activation.quants(), activation.scales(), output, batchSize, rows, cols); - return output.toHeapArray(); - } - - private static float[] referenceQ4K( - byte[] weights, float[] input, int batchSize, int rows, int cols) { - float[] output = new float[batchSize * rows]; - VectorUtil.ggufQ4_KQ8_KBatchedMatmul( - input, - MemorySegment.ofArray(weights), - batchSize, - rows, - cols, - output, - new byte[batchSize * cols], - new float[batchSize * cols / SUPER_BLOCK], - new short[batchSize * cols / SUM_BLOCK]); - return output; - } - - private static float[] referenceQ6K( - byte[] weights, float[] input, int batchSize, int rows, int cols) { - float[] output = new float[batchSize * rows]; - VectorUtil.ggufQ6_KQ8_KBatchedMatmul( - input, - MemorySegment.ofArray(weights), - batchSize, - rows, - cols, - output, - new byte[batchSize * cols], - new float[batchSize * cols / SUPER_BLOCK]); - return output; - } - - /** - * Asserts the guaranteed numeric contract: the kernel differs from the production CPU kernel by - * at most two float roundings per super-block, scaled to the magnitude of the reference result. - */ - private static void assertWithinFusedMultiplyAddBudget( - float[] actual, float[] expected, int cols) { - assertThat(actual).hasSameSizeAs(expected); - float magnitude = 0.0f; - for (float value : expected) { - magnitude = Math.max(magnitude, Math.abs(value)); - } - float tolerance = 2.0f * (cols / SUPER_BLOCK) * Math.ulp(magnitude); - for (int index = 0; index < expected.length; index++) { - assertThat(actual[index]) - .as("element %d within %s of %s", index, tolerance, expected[index]) - .isCloseTo(expected[index], within(tolerance)); - } - } - - /** Pseudo-random Q4_K super-blocks spanning the whole six-bit scale and minimum range. */ - private static byte[] randomQ4KMatrix(int rows, int cols, long seed) { - Random random = new Random(seed); - int blocks = rows * cols / SUPER_BLOCK; - byte[] weights = new byte[blocks * KQuantProjectionKernel.Q4_K_BLOCK_BYTES]; - int[] scales = new int[8]; - int[] minimums = new int[8]; - for (int block = 0; block < blocks; block++) { - for (int group = 0; group < 8; group++) { - scales[group] = random.nextInt(64); - minimums[group] = random.nextInt(64); - } - writeQ4KBlock( - weights, - block, - Float.floatToFloat16(0.002f + random.nextFloat() * 0.02f), - Float.floatToFloat16(0.001f + random.nextFloat() * 0.01f), - scales, - minimums, - random); - } - return weights; - } - - /** - * Q4_K super-blocks whose contribution is exactly representable: the two scales are powers of two - * and the six-bit group scales and minimums stay small enough that no partial sum leaves the - * exactly representable integer range of a float. - */ - private static byte[] exactQ4KMatrix(int rows, int cols, long seed) { - Random random = new Random(seed); - int blocks = rows * cols / SUPER_BLOCK; - byte[] weights = new byte[blocks * KQuantProjectionKernel.Q4_K_BLOCK_BYTES]; - int[] scales = new int[8]; - int[] minimums = new int[8]; - for (int block = 0; block < blocks; block++) { - for (int group = 0; group < 8; group++) { - scales[group] = 1 + random.nextInt(3); - minimums[group] = random.nextInt(4); - } - writeQ4KBlock( - weights, - block, - Float.floatToFloat16(1.0f), - Float.floatToFloat16(0.25f), - scales, - minimums, - random); - } - return weights; - } - - private static void writeQ4KBlock( - byte[] weights, - int block, - short delta, - short deltaMin, - int[] scales, - int[] minimums, - Random random) { - int offset = block * KQuantProjectionKernel.Q4_K_BLOCK_BYTES; - weights[offset] = (byte) delta; - weights[offset + 1] = (byte) (delta >>> 8); - weights[offset + 2] = (byte) deltaMin; - weights[offset + 3] = (byte) (deltaMin >>> 8); - for (int group = 0; group < 4; group++) { - weights[offset + 4 + group] = - (byte) ((scales[group] & 0x3F) | ((scales[group + 4] >>> 4) << 6)); - weights[offset + 8 + group] = - (byte) ((minimums[group] & 0x3F) | ((minimums[group + 4] >>> 4) << 6)); - weights[offset + 12 + group] = - (byte) ((scales[group + 4] & 0x0F) | ((minimums[group + 4] & 0x0F) << 4)); - } - for (int index = 0; index < 128; index++) { - weights[offset + 16 + index] = (byte) random.nextInt(256); - } - } - - /** Pseudo-random Q6_K super-blocks spanning the whole signed int8 group-scale range. */ - private static byte[] randomQ6KMatrix(int rows, int cols, long seed) { - Random random = new Random(seed); - int blocks = rows * cols / SUPER_BLOCK; - byte[] weights = new byte[blocks * KQuantProjectionKernel.Q6_K_BLOCK_BYTES]; - for (int block = 0; block < blocks; block++) { - int offset = block * KQuantProjectionKernel.Q6_K_BLOCK_BYTES; - for (int index = 0; index < 192; index++) { - weights[offset + index] = (byte) random.nextInt(256); - } - for (int index = 0; index < 16; index++) { - weights[offset + 192 + index] = (byte) (random.nextInt(65) - 32); - } - short delta = Float.floatToFloat16(0.002f + random.nextFloat() * 0.02f); - weights[offset + 208] = (byte) delta; - weights[offset + 209] = (byte) (delta >>> 8); - } - return weights; - } - - /** Q6_K super-blocks with a power-of-two delta and small group scales. */ - private static byte[] exactQ6KMatrix(int rows, int cols, long seed) { - Random random = new Random(seed); - int blocks = rows * cols / SUPER_BLOCK; - byte[] weights = new byte[blocks * KQuantProjectionKernel.Q6_K_BLOCK_BYTES]; - for (int block = 0; block < blocks; block++) { - int offset = block * KQuantProjectionKernel.Q6_K_BLOCK_BYTES; - for (int index = 0; index < 192; index++) { - weights[offset + index] = (byte) random.nextInt(256); - } - for (int index = 0; index < 16; index++) { - weights[offset + 192 + index] = (byte) (random.nextInt(5) - 2); - } - short delta = Float.floatToFloat16(0.5f); - weights[offset + 208] = (byte) delta; - weights[offset + 209] = (byte) (delta >>> 8); - } - return weights; - } - - private static float[] randomFloats(int length, long seed) { - Random random = new Random(seed); - float[] values = new float[length]; - for (int index = 0; index < length; index++) { - values[index] = random.nextFloat(-2.0f, 2.0f); - } - return values; - } - - /** - * Activations that quantize to Q8_K without rounding error: every value is a small integer times - * 2^-6, and each super-block's extremum is exactly -127 * 2^-6, so the Q8_K scale is 2^-6 and - * each stored quant is the integer it started from. - */ - private static float[] exactFloats(int batchSize, int cols, long seed) { - Random random = new Random(seed); - float[] values = new float[batchSize * cols]; - float step = 1.0f / 64.0f; - for (int batch = 0; batch < batchSize; batch++) { - for (int block = 0; block < cols / SUPER_BLOCK; block++) { - int offset = batch * cols + block * SUPER_BLOCK; - values[offset] = -127.0f * step; - for (int index = 1; index < SUPER_BLOCK; index++) { - values[offset + index] = (random.nextInt(7) - 3) * step; - } - } - } - return values; - } -} diff --git a/backend-tornado/src/test/java/com/integrallis/models/backend/tornado/LargeModelEligibilityTest.java b/backend-tornado/src/test/java/com/integrallis/models/backend/tornado/LargeModelEligibilityTest.java deleted file mode 100644 index be15d7754..000000000 --- a/backend-tornado/src/test/java/com/integrallis/models/backend/tornado/LargeModelEligibilityTest.java +++ /dev/null @@ -1,292 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.backend.tornado; - -import static org.assertj.core.api.Assertions.assertThat; - -import java.util.List; -import java.util.stream.Stream; -import org.junit.jupiter.api.Test; -import org.junit.jupiter.params.ParameterizedTest; -import org.junit.jupiter.params.provider.Arguments; -import org.junit.jupiter.params.provider.MethodSource; - -/** - * Device-independent capacity arithmetic for 27B-class Q4_K_M models. - * - *

Every model constant below is read from this repository, not assumed: - * - *

    - *
  • weights {@code 1_650_027_640 + 15_130_165_248} are the resident and routed-expert byte - * counts asserted in {@code Gemma4LargeModelFixtureSlowTest} for the pinned {@code - * gemma-4-26B-A4B-it-Q4_K_M.gguf}; - *
  • the KV geometry (30 layers, 25 sliding with keyDim 2048, 5 full with keyDim 1024, ring - * capacity {@code slidingWindow 1024 + prefillCapacity 1}) is asserted in {@code - * Gemma4ConfigTest.parsesThePinnedGemma426BA4BMetadata} and built in {@code - * Gemma4KvCache.create}; - *
  • the retained plan count is the number of distinct {@code (weight, batch shape)} keys {@code - * TornadoGgufBatchedMatrixKernel} would create for this graph. - *
- * - *

Device global-memory figures are driver-reported values, which sit below the nameplate - * capacity: the A40-4Q gate reported 3,917.7 MiB on a 4,096 MiB profile. - */ -class LargeModelEligibilityTest { - - private static final long MIB = 1024L * 1024L; - private static final long GIB = 1024L * MIB; - - /** Gemma 4 26B-A4B IT Q4_K_M: resident tensors plus routed experts. */ - private static final long GEMMA4_WEIGHT_BYTES = 1_650_027_640L + 15_130_165_248L; - - /** - * Distinct retained plans: 30 layers x (128 experts x 2 tensors + 4 resident projections) x 2 - * retained batch shapes. - */ - private static final int GEMMA4_PLAN_COUNT = 30 * (128 * 2 + 4) * 2; - - /** Device activation, scale, and output buffers summed over every retained plan. */ - private static final long GEMMA4_PLAN_SCRATCH_BYTES = 2_733_284_400L; - - /** Largest single tensor: the 262,144 x 2,816 tied embedding at Q8_0. */ - private static final long GEMMA4_LARGEST_ALLOCATION_BYTES = 262_144L * 2_816L / 32L * 34L; - - private static final long QWEN3_06B_Q4_0_BYTES = 428_970_080L; - - /** KV bytes the {@code LayeredKvCache} would hold at a runtime context length. */ - private static long gemma4KvBytes(int contextLength) { - long slidingRing = 25L * 2L * 2_048L * 4L * (1_024L + 1L); - long full = 5L * 2L * 1_024L * 4L * Math.min(contextLength, 262_144L); - return slidingRing + full; - } - - private static DeviceMemoryRequest gemma4(int contextLength, int planCount, long planScratch) { - return DeviceMemoryRequest.detailed("gemma-4-26B-A4B-it-Q4_K_M.gguf") - .weightBytes(GEMMA4_WEIGHT_BYTES) - .retainedShapes(2) - .retainedPlanCount(planCount) - .planScratchBytes(planScratch) - .largestAllocationBytes(GEMMA4_LARGEST_ALLOCATION_BYTES) - .deviceKvCacheBytes(gemma4KvBytes(contextLength)) - .build(); - } - - private static AcceleratorEligibility.DeviceCapabilities gpu(String name, long globalBytes) { - return new AcceleratorEligibility.DeviceCapabilities(name, "PTX", "GPU", globalBytes, 8L * GIB); - } - - static Stream deviceCapacityTable() { - // device, driver-reported global memory, context length, expected outcome marker - return Stream.of( - Arguments.of("A40 24 GB", 23L * GIB, 4_096, "short by"), - Arguments.of("A40 24 GB", 23L * GIB, 32_768, "short by"), - Arguments.of("A40 24 GB", 23L * GIB, 262_144, "short by"), - Arguments.of("L40S 48 GB", 44L * GIB, 4_096, "SHARED_WEIGHT_UPLOAD"), - Arguments.of("L40S 48 GB", 44L * GIB, 32_768, "SHARED_WEIGHT_UPLOAD"), - Arguments.of("L40S 48 GB", 44L * GIB, 262_144, "SHARED_WEIGHT_UPLOAD"), - Arguments.of("H100 80 GB", 79L * GIB, 4_096, "eager readiness"), - Arguments.of("H100 80 GB", 79L * GIB, 262_144, "eager readiness")); - } - - @ParameterizedTest(name = "{0} at context {2} is refused because of \"{3}\"") - @MethodSource("deviceCapacityTable") - void refusesTheShippedPlanShapeForA27bClassModelOnEveryDeviceSize( - String deviceName, long globalBytes, int contextLength, String expectedReason) { - AcceleratorEligibility.Decision decision = - AcceleratorEligibility.select( - List.of(gpu(deviceName, globalBytes)), - gemma4(contextLength, GEMMA4_PLAN_COUNT, GEMMA4_PLAN_SCRATCH_BYTES)); - - assertThat(decision.eligible()).isFalse(); - assertThat(decision.reason()).contains(expectedReason); - assertThat(decision.reason()).contains(deviceName); - } - - @Test - void itemisesEveryTermOfTheRefusedBudget() { - DeviceMemoryRequest request = gemma4(4_096, GEMMA4_PLAN_COUNT, GEMMA4_PLAN_SCRATCH_BYTES); - DeviceBudget budget = - AcceleratorEligibility.budget(request, PlanShapeStrategy.PER_SHAPE_WHOLE_MODEL); - - assertThat(budget.deviceWeightBytes()).isEqualTo(2 * GEMMA4_WEIGHT_BYTES); - assertThat(budget.kvCacheBytes()).isEqualTo(587_612_160L); - assertThat(budget.planScratchBytes()).isEqualTo(GEMMA4_PLAN_SCRATCH_BYTES); - assertThat(budget.baseOverheadBytes()).isEqualTo(256 * MIB); - assertThat(budget.totalBytes()).isEqualTo(37_149_717_792L); - assertThat(budget.hostWeightCopyBytes()).isEqualTo(2 * GEMMA4_WEIGHT_BYTES); - assertThat(budget.estimatedPlanCompileTime().toSeconds()).isEqualTo(948L); - } - - @Test - void sharedWeightUploadHalvesTheWeightTermButIsNotImplemented() { - DeviceMemoryRequest request = gemma4(4_096, GEMMA4_PLAN_COUNT, GEMMA4_PLAN_SCRATCH_BYTES); - DeviceBudget shared = - AcceleratorEligibility.budget(request, PlanShapeStrategy.SHARED_WEIGHT_UPLOAD); - - assertThat(shared.deviceWeightBytes()).isEqualTo(GEMMA4_WEIGHT_BYTES); - assertThat(shared.totalBytes()).isEqualTo(20_369_524_904L); - assertThat(PlanShapeStrategy.SHARED_WEIGHT_UPLOAD.implemented()).isFalse(); - assertThat(PlanShapeStrategy.SHARED_WEIGHT_UPLOAD.limitation()).contains("ByteArray"); - assertThat(PlanShapeStrategy.PER_SHAPE_WHOLE_MODEL.implemented()).isTrue(); - assertThat(PlanShapeStrategy.PER_SHAPE_WHOLE_MODEL.limitation()).isNull(); - assertThat(PlanShapeStrategy.BOUNDED_RESIDENT_WORKING_SET.capacityComputable()).isFalse(); - } - - @Test - void admitsAnEightyGigabyteDeviceOnceThePerExpertPlansAreGone() { - int residentOnlyPlans = 30 * 4 * 2; - AcceleratorEligibility.Decision decision = - AcceleratorEligibility.select( - List.of(gpu("H100 80 GB", 79L * GIB)), gemma4(4_096, residentOnlyPlans, 90_302_490L)); - - assertThat(decision.eligible()).isTrue(); - assertThat(decision.strategy()).isEqualTo(PlanShapeStrategy.PER_SHAPE_WHOLE_MODEL); - assertThat(decision.budget().estimatedPlanCompileTime().toSeconds()).isEqualTo(14L); - } - - @Test - void refusesAFileSizeOnlyBudgetForALargeModel() { - AcceleratorEligibility.Decision decision = - AcceleratorEligibility.select( - List.of(gpu("H100 80 GB", 79L * GIB)), - DeviceMemoryRequest.ofModelFile( - "gemma-4-26B-A4B-it-Q4_K_M.gguf", 16_796_015_136L, true)); - - assertThat(decision.eligible()).isFalse(); - assertThat(decision.reason()).contains("file-size-only device budget"); - assertThat(decision.reason()).contains("15.64 GiB of weights"); - assertThat(decision.reason()).contains("8.00 GiB"); - assertThat(decision.reason()).contains("DeviceMemoryRequest.detailed"); - } - - @Test - void refusesATensorThatCannotFitInATornadoByteArray() { - AcceleratorEligibility.Decision decision = - AcceleratorEligibility.select( - List.of(gpu("H100 80 GB", 79L * GIB)), - DeviceMemoryRequest.detailed("oversized.gguf") - .weightBytes(10L * GIB) - .retainedShapes(1) - .retainedPlanCount(8) - .planScratchBytes(MIB) - .largestAllocationBytes(3L * GIB) - .build()); - - assertThat(decision.eligible()).isFalse(); - assertThat(decision.reason()).contains("largest device allocation is 3.00 GiB"); - assertThat(decision.reason()).contains("2.00 GiB"); - assertThat(decision.reason()).contains("must be split"); - } - - @Test - void refusesADeviceWhoseSingleAllocationLimitIsTooSmall() { - AcceleratorEligibility.Decision decision = - AcceleratorEligibility.select( - List.of( - new AcceleratorEligibility.DeviceCapabilities( - "A40-4Q", "PTX", "GPU", 79L * GIB, 512L * MIB)), - gemma4(4_096, 240, 90_302_490L)); - - assertThat(decision.eligible()).isFalse(); - assertThat(decision.reason()).contains("single allocation of at most 0.50 GiB"); - assertThat(decision.reason()).contains("needs one of 0.73 GiB"); - } - - @Test - void namesEveryDiscoveredDeviceWhenNoneQualifies() { - AcceleratorEligibility.Decision decision = - AcceleratorEligibility.select( - List.of( - new AcceleratorEligibility.DeviceCapabilities( - "AMD Radeon", "OPENCL", "GPU", 24L * GIB, 4L * GIB), - new AcceleratorEligibility.DeviceCapabilities( - "host CPU", "JAVA", "CPU", 64L * GIB, 16L * GIB)), - gemma4(4_096, 240, 90_302_490L)); - - assertThat(decision.eligible()).isFalse(); - assertThat(decision.reason()).contains("PTX or CUDA backend on a GPU device"); - assertThat(decision.reason()).contains("AMD Radeon [OPENCL/GPU]"); - assertThat(decision.reason()).contains("host CPU [JAVA/CPU]"); - } - - @Test - void reportsNoDevicesWhenTheRuntimeFoundNone() { - AcceleratorEligibility.Decision decision = - AcceleratorEligibility.select(List.of(), gemma4(4_096, 240, 90_302_490L)); - - assertThat(decision.eligible()).isFalse(); - assertThat(decision.reason()).contains("no devices"); - } - - @Test - void keepsThePublishedCoarseBudgetForTheQualifiedSmallModel() { - DeviceMemoryRequest request = - DeviceMemoryRequest.ofModelFile("Qwen3-0.6B-Q4_0.gguf", QWEN3_06B_Q4_0_BYTES, true); - DeviceBudget budget = - AcceleratorEligibility.budget(request, PlanShapeStrategy.PER_SHAPE_WHOLE_MODEL); - - assertThat(request.detailed()).isFalse(); - assertThat(budget.totalBytes()).isEqualTo(1_126_375_616L); - assertThat(budget.kvCacheBytes()).isZero(); - assertThat(budget.estimatedPlanCompileTime()).isZero(); - } - - @Test - void carriesADeviceResidentKvCacheOntoAnExistingRequest() { - DeviceMemoryRequest base = - DeviceMemoryRequest.ofModelFile("Qwen3-0.6B-Q4_0.gguf", QWEN3_06B_Q4_0_BYTES, false); - - DeviceMemoryRequest withKv = base.withDeviceKvCacheBytes(64 * MIB); - - assertThat(base.deviceKvCacheBytes()).isZero(); - assertThat(withKv.deviceKvCacheBytes()).isEqualTo(64 * MIB); - assertThat(withKv.retainedShapes()).isEqualTo(1); - assertThat( - AcceleratorEligibility.budget(withKv, PlanShapeStrategy.PER_SHAPE_WHOLE_MODEL) - .totalBytes()) - .isEqualTo(QWEN3_06B_Q4_0_BYTES + 64 * MIB + 256 * MIB); - } - - @Test - void rejectsMalformedRequests() { - assertThat( - org.assertj.core.api.Assertions.catchThrowable( - () -> DeviceMemoryRequest.ofModelFile(" ", 1, true))) - .isInstanceOf(IllegalArgumentException.class); - assertThat( - org.assertj.core.api.Assertions.catchThrowable( - () -> DeviceMemoryRequest.detailed("m").weightBytes(0).build())) - .isInstanceOf(IllegalArgumentException.class); - assertThat( - org.assertj.core.api.Assertions.catchThrowable( - () -> DeviceMemoryRequest.detailed("m").weightBytes(1).retainedShapes(0).build())) - .isInstanceOf(IllegalArgumentException.class); - assertThat( - org.assertj.core.api.Assertions.catchThrowable( - () -> - DeviceMemoryRequest.detailed("m") - .weightBytes(1) - .deviceKvCacheBytes(-1) - .build())) - .isInstanceOf(IllegalArgumentException.class); - assertThat( - org.assertj.core.api.Assertions.catchThrowable( - () -> - PlanShapeStrategy.BOUNDED_RESIDENT_WORKING_SET.deviceWeightBytes( - DeviceMemoryRequest.ofModelFile("m", 1, false)))) - .isInstanceOf(IllegalStateException.class); - } -} diff --git a/backend-tornado/src/test/java/com/integrallis/models/backend/tornado/Q4ProjectionKernelTest.java b/backend-tornado/src/test/java/com/integrallis/models/backend/tornado/Q4ProjectionKernelTest.java deleted file mode 100644 index dff73fc14..000000000 --- a/backend-tornado/src/test/java/com/integrallis/models/backend/tornado/Q4ProjectionKernelTest.java +++ /dev/null @@ -1,200 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.backend.tornado; - -import static org.assertj.core.api.Assertions.assertThat; -import static org.assertj.core.api.Assertions.assertThatThrownBy; -import static org.assertj.core.api.Assertions.within; - -import com.integrallis.vectors.core.VectorUtil; -import java.lang.foreign.MemorySegment; -import java.util.Random; -import org.junit.jupiter.api.Test; -import uk.ac.manchester.tornado.api.types.arrays.ByteArray; -import uk.ac.manchester.tornado.api.types.arrays.FloatArray; - -class Q4ProjectionKernelTest { - - @Test - void matchesProductionQ4ByQ8ProjectionAcrossBatches() { - int batchSize = 3; - int rows = 17; - int cols = 96; - byte[] weights = randomQ4Matrix(rows, cols, 31L); - float[] input = randomFloats(batchSize * cols, 37L); - byte[] activations = new byte[batchSize * cols]; - float[] activationScales = new float[batchSize * cols / 32]; - Q4ProjectionKernel.quantize(input, activations, activationScales, batchSize, cols); - float[] expected = vectorApiProjection(weights, input, batchSize, rows, cols); - ByteArray deviceWeights = ByteArray.fromArray(weights); - ByteArray deviceActivations = ByteArray.fromArray(activations); - FloatArray deviceScales = FloatArray.fromArray(activationScales); - FloatArray deviceOutput = new FloatArray(batchSize * rows); - - Q4ProjectionKernel.multiply( - deviceWeights, deviceActivations, deviceScales, deviceOutput, batchSize, rows, cols); - - float[] actual = deviceOutput.toHeapArray(); - assertThat(actual).hasSameSizeAs(expected); - for (int index = 0; index < expected.length; index++) { - assertThat(actual[index]).isCloseTo(expected[index], within(2.0e-5f)); - } - } - - @Test - void matchesTwoProductionProjectionsWithOnePreparedActivation() { - int batchSize = 3; - int firstRows = 11; - int secondRows = 7; - int cols = 96; - byte[] firstWeights = randomQ4Matrix(firstRows, cols, 41L); - byte[] secondWeights = randomQ4Matrix(secondRows, cols, 43L); - float[] input = randomFloats(batchSize * cols, 47L); - byte[] activations = new byte[batchSize * cols]; - float[] scales = new float[batchSize * cols / 32]; - Q4ProjectionKernel.quantize(input, activations, scales, batchSize, cols); - float[] expectedFirst = vectorApiProjection(firstWeights, input, batchSize, firstRows, cols); - float[] expectedSecond = vectorApiProjection(secondWeights, input, batchSize, secondRows, cols); - FloatArray actualFirst = new FloatArray(batchSize * firstRows); - FloatArray actualSecond = new FloatArray(batchSize * secondRows); - - Q4ProjectionKernel.multiplyDual( - ByteArray.fromArray(firstWeights), - firstRows, - ByteArray.fromArray(secondWeights), - secondRows, - ByteArray.fromArray(activations), - FloatArray.fromArray(scales), - actualFirst, - actualSecond, - batchSize, - cols); - - assertClose(actualFirst.toHeapArray(), expectedFirst); - assertClose(actualSecond.toHeapArray(), expectedSecond); - } - - @Test - void matchesThreeProductionProjectionsWithOnePreparedActivation() { - int batchSize = 2; - int firstRows = 9; - int secondRows = 7; - int thirdRows = 5; - int cols = 64; - byte[] firstWeights = randomQ4Matrix(firstRows, cols, 53L); - byte[] secondWeights = randomQ4Matrix(secondRows, cols, 59L); - byte[] thirdWeights = randomQ4Matrix(thirdRows, cols, 61L); - float[] input = randomFloats(batchSize * cols, 67L); - byte[] activations = new byte[batchSize * cols]; - float[] scales = new float[batchSize * cols / 32]; - Q4ProjectionKernel.quantize(input, activations, scales, batchSize, cols); - float[] expectedFirst = vectorApiProjection(firstWeights, input, batchSize, firstRows, cols); - float[] expectedSecond = vectorApiProjection(secondWeights, input, batchSize, secondRows, cols); - float[] expectedThird = vectorApiProjection(thirdWeights, input, batchSize, thirdRows, cols); - FloatArray actualFirst = new FloatArray(batchSize * firstRows); - FloatArray actualSecond = new FloatArray(batchSize * secondRows); - FloatArray actualThird = new FloatArray(batchSize * thirdRows); - - Q4ProjectionKernel.multiplyTriple( - ByteArray.fromArray(firstWeights), - firstRows, - ByteArray.fromArray(secondWeights), - secondRows, - ByteArray.fromArray(thirdWeights), - thirdRows, - ByteArray.fromArray(activations), - FloatArray.fromArray(scales), - actualFirst, - actualSecond, - actualThird, - batchSize, - cols); - - assertClose(actualFirst.toHeapArray(), expectedFirst); - assertClose(actualSecond.toHeapArray(), expectedSecond); - assertClose(actualThird.toHeapArray(), expectedThird); - } - - @Test - void rejectsMismatchedActivationScaleStorage() { - assertThatThrownBy( - () -> - Q4ProjectionKernel.validate( - new ByteArray(18), - new ByteArray(32), - new FloatArray(0), - new FloatArray(1), - 1, - 1, - 32)) - .isInstanceOf(IllegalArgumentException.class) - .hasMessageContaining("scales"); - } - - private static float[] vectorApiProjection( - byte[] weights, float[] input, int batchSize, int rows, int cols) { - float[] output = new float[batchSize * rows]; - MemorySegment weightSegment = MemorySegment.ofArray(weights); - for (int batch = 0; batch < batchSize; batch++) { - float[] query = new float[cols]; - float[] projected = new float[rows]; - System.arraycopy(input, batch * cols, query, 0, cols); - VectorUtil.ggufQ4_0Q8_0BatchDotProduct( - query, - weightSegment, - rows, - cols, - projected, - new byte[cols], - new float[cols / 32], - new int[(cols + 3) / 4]); - System.arraycopy(projected, 0, output, batch * rows, rows); - } - return output; - } - - private static void assertClose(float[] actual, float[] expected) { - assertThat(actual).hasSameSizeAs(expected); - for (int index = 0; index < expected.length; index++) { - assertThat(actual[index]).isCloseTo(expected[index], within(2.0e-5f)); - } - } - - private static byte[] randomQ4Matrix(int rows, int cols, long seed) { - Random random = new Random(seed); - int blocks = rows * cols / 32; - byte[] weights = new byte[blocks * 18]; - for (int block = 0; block < blocks; block++) { - short scale = Float.floatToFloat16(0.001f + random.nextFloat() * 0.05f); - int offset = block * 18; - weights[offset] = (byte) scale; - weights[offset + 1] = (byte) (scale >>> 8); - for (int quant = 0; quant < 16; quant++) { - weights[offset + 2 + quant] = (byte) random.nextInt(256); - } - } - return weights; - } - - private static float[] randomFloats(int length, long seed) { - Random random = new Random(seed); - float[] values = new float[length]; - for (int index = 0; index < length; index++) { - values[index] = random.nextFloat(-2.0f, 2.0f); - } - return values; - } -} diff --git a/backend-tornado/src/test/java/com/integrallis/models/backend/tornado/TornadoBackendIntegrationTest.java b/backend-tornado/src/test/java/com/integrallis/models/backend/tornado/TornadoBackendIntegrationTest.java deleted file mode 100644 index e8884ecbf..000000000 --- a/backend-tornado/src/test/java/com/integrallis/models/backend/tornado/TornadoBackendIntegrationTest.java +++ /dev/null @@ -1,59 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.backend.tornado; - -import static org.assertj.core.api.Assertions.assertThat; -import static org.junit.jupiter.api.Assumptions.assumeTrue; - -import java.nio.file.Files; -import java.nio.file.Path; -import org.junit.jupiter.api.Tag; -import org.junit.jupiter.api.Test; - -@Tag("integration") -class TornadoBackendIntegrationTest { - - @Test - void loadsAndRunsThePinnedQwenModelThroughAutomaticSelectionWhenProvided() { - String configured = System.getProperty("models.fixtures.qwen306BQ40", ""); - assumeTrue(!configured.isBlank(), "set -Dmodels.fixtures.qwen306BQ40="); - Path model = Path.of(configured).toAbsolutePath().normalize(); - assumeTrue(Files.isRegularFile(model), "Qwen fixture is not installed"); - boolean required = Boolean.getBoolean("models.accelerator.required"); - TornadoBackendOptions options = new TornadoBackendOptions(true, true, required, 32); - - try (TornadoBackendRuntime runtime = TornadoBackend.open(model, options)) { - int token = runtime.backend().tokenizer().bosToken(); - if (token < 0 || token >= runtime.backend().metadata().vocabSize()) { - token = 0; - } - float[] logits = runtime.backend().prefill(new int[] {token}, 0); - - assertThat(logits).hasSize(runtime.backend().metadata().vocabSize()); - for (float logit : logits) { - assertThat(Float.isFinite(logit)).isTrue(); - } - if (required) { - assertThat(runtime.status().accelerated()).isTrue(); - assertThat(runtime.status().readinessTime()).isPositive(); - } - String expected = System.getProperty("models.accelerator.expected", ""); - if (!expected.isBlank()) { - assertThat(runtime.status().accelerated()).isEqualTo(Boolean.parseBoolean(expected)); - } - } - } -} diff --git a/backend-tornado/src/test/java/com/integrallis/models/backend/tornado/TornadoBackendOptionsTest.java b/backend-tornado/src/test/java/com/integrallis/models/backend/tornado/TornadoBackendOptionsTest.java deleted file mode 100644 index 8146700ce..000000000 --- a/backend-tornado/src/test/java/com/integrallis/models/backend/tornado/TornadoBackendOptionsTest.java +++ /dev/null @@ -1,40 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.backend.tornado; - -import static org.assertj.core.api.Assertions.assertThat; - -import org.junit.jupiter.api.Test; - -class TornadoBackendOptionsTest { - - @Test - void defaultsPrepareBothPrefillAndDecodeAndAllowCpuFallback() { - TornadoBackendOptions options = TornadoBackendOptions.defaults(); - - assertThat(options.accelerateDecode()).isTrue(); - assertThat(options.eagerReadiness()).isTrue(); - assertThat(options.requireAccelerator()).isFalse(); - assertThat(options.executionBatchSize()).isEqualTo(32); - } - - @Test - void validatesTheFixedDeviceBatchShape() { - org.assertj.core.api.Assertions.assertThatIllegalArgumentException() - .isThrownBy(() -> new TornadoBackendOptions(true, true, false, 3)) - .withMessageContaining("executionBatchSize"); - } -} diff --git a/backend-tornado/src/test/java/com/integrallis/models/backend/tornado/TornadoBackendStatusTest.java b/backend-tornado/src/test/java/com/integrallis/models/backend/tornado/TornadoBackendStatusTest.java deleted file mode 100644 index 6072cb0b7..000000000 --- a/backend-tornado/src/test/java/com/integrallis/models/backend/tornado/TornadoBackendStatusTest.java +++ /dev/null @@ -1,55 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.backend.tornado; - -import static org.assertj.core.api.Assertions.assertThat; -import static org.assertj.core.api.Assertions.assertThatThrownBy; - -import java.time.Duration; -import org.junit.jupiter.api.Test; - -class TornadoBackendStatusTest { - - @Test - void normalizesAValidStatus() { - TornadoBackendStatus status = - new TornadoBackendStatus( - true, " NVIDIA A40 ", " eligible ", 1024, Duration.ofSeconds(2)); - - assertThat(status.device()).isEqualTo("NVIDIA A40"); - assertThat(status.reason()).isEqualTo("eligible"); - assertThat(status.requiredDeviceBytes()).isEqualTo(1024); - assertThat(status.readinessTime()).isEqualTo(Duration.ofSeconds(2)); - } - - @Test - void rejectsInvalidStatusValues() { - assertThatThrownBy(() -> new TornadoBackendStatus(false, " ", "unavailable", 0, Duration.ZERO)) - .isInstanceOf(IllegalArgumentException.class) - .hasMessageContaining("device"); - assertThatThrownBy(() -> new TornadoBackendStatus(false, "CPU", " ", 0, Duration.ZERO)) - .isInstanceOf(IllegalArgumentException.class) - .hasMessageContaining("reason"); - assertThatThrownBy( - () -> new TornadoBackendStatus(false, "CPU", "unavailable", -1, Duration.ZERO)) - .isInstanceOf(IllegalArgumentException.class) - .hasMessageContaining("requiredDeviceBytes"); - assertThatThrownBy( - () -> new TornadoBackendStatus(false, "CPU", "unavailable", 0, Duration.ofNanos(-1))) - .isInstanceOf(IllegalArgumentException.class) - .hasMessageContaining("readinessTime"); - } -} diff --git a/backend-tornado/src/test/java/com/integrallis/models/backend/tornado/TornadoBackendTest.java b/backend-tornado/src/test/java/com/integrallis/models/backend/tornado/TornadoBackendTest.java deleted file mode 100644 index f7b185a12..000000000 --- a/backend-tornado/src/test/java/com/integrallis/models/backend/tornado/TornadoBackendTest.java +++ /dev/null @@ -1,50 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.backend.tornado; - -import static org.assertj.core.api.Assertions.assertThatIllegalArgumentException; -import static org.assertj.core.api.Assertions.assertThatNullPointerException; - -import com.integrallis.models.api.BackendConfiguration; -import java.nio.file.Files; -import java.nio.file.Path; -import org.junit.jupiter.api.Test; -import org.junit.jupiter.api.io.TempDir; - -class TornadoBackendTest { - - @Test - void rejectsAMissingModelBeforeInspectingAcceleratorDrivers(@TempDir Path directory) { - Path missing = directory.resolve("missing.gguf"); - - assertThatIllegalArgumentException() - .isThrownBy(() -> TornadoBackend.open(missing)) - .withMessageContaining("regular model file"); - } - - @Test - void validatesConfigurationBeforeInspectingAcceleratorDrivers(@TempDir Path directory) - throws Exception { - Path placeholder = Files.createFile(directory.resolve("model.gguf")); - - assertThatNullPointerException() - .isThrownBy(() -> TornadoBackend.open(placeholder, null, TornadoBackendOptions.defaults())) - .withMessage("backendConfiguration"); - assertThatNullPointerException() - .isThrownBy(() -> TornadoBackend.open(placeholder, BackendConfiguration.empty(), null)) - .withMessage("options"); - } -} diff --git a/backend-tornado/src/test/java/com/integrallis/models/backend/tornado/TornadoGgufBatchedMatrixKernelTest.java b/backend-tornado/src/test/java/com/integrallis/models/backend/tornado/TornadoGgufBatchedMatrixKernelTest.java deleted file mode 100644 index 2c4d3ecf5..000000000 --- a/backend-tornado/src/test/java/com/integrallis/models/backend/tornado/TornadoGgufBatchedMatrixKernelTest.java +++ /dev/null @@ -1,210 +0,0 @@ -/* - * Copyright 2025-2026 Integrallis Software, LLC - * - * Licensed under the Apache License, Version 2.0 (the "License"); - * you may not use this file except in compliance with the License. - * You may obtain a copy of the License at - * - * https://www.apache.org/licenses/LICENSE-2.0 - * - * Unless required by applicable law or agreed to in writing, software - * distributed under the License is distributed on an "AS IS" BASIS, - * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. - * See the License for the specific language governing permissions and - * limitations under the License. - */ -package com.integrallis.models.backend.tornado; - -import static org.assertj.core.api.Assertions.assertThat; -import static org.assertj.core.api.Assertions.assertThatThrownBy; - -import com.integrallis.models.backend.purejava.gguf.GgufTensorType; -import com.integrallis.models.backend.purejava.plan.PureJavaPlanConfiguration; -import java.lang.foreign.MemorySegment; -import org.junit.jupiter.api.Test; - -class TornadoGgufBatchedMatrixKernelTest { - - @Test - void acceleratesQ4PrefillButLeavesDecodeAndUnsupportedFormatsOnTheJavaPath() { - try (TornadoGgufBatchedMatrixKernel kernel = new TornadoGgufBatchedMatrixKernel()) { - assertThat(kernel.executionBatchSize()).isEqualTo(32); - assertThat(kernel.supports(GgufTensorType.Q4_0)).isTrue(); - assertThat(kernel.supports(GgufTensorType.Q4_K)).isTrue(); - assertThat(kernel.supports(GgufTensorType.Q6_K)).isTrue(); - assertThat(kernel.supports(GgufTensorType.Q5_K)).isFalse(); - assertThat(kernel.supports(GgufTensorType.Q8_0)).isFalse(); - assertThat(kernel.supports(GgufTensorType.F32)).isFalse(); - assertThat(kernel.isEligible(GgufTensorType.Q4_0, 1, 3072, 1024)).isFalse(); - assertThat(kernel.isEligible(GgufTensorType.Q4_0, 4, 3072, 1024)).isTrue(); - assertThat(kernel.isEligible(GgufTensorType.Q4_0, 30, 3072, 1024)).isTrue(); - assertThat(kernel.isEligible(GgufTensorType.Q4_0, 32, 3072, 1024)).isTrue(); - assertThat(kernel.isEligible(GgufTensorType.Q4_0, 33, 3072, 1024)).isFalse(); - assertThat(kernel.isEligible(GgufTensorType.Q4_0, 4, 32, 32)).isFalse(); - assertThat(kernel.supportsDual(GgufTensorType.Q4_0, GgufTensorType.Q4_0)).isTrue(); - assertThat( - kernel.supportsTriple(GgufTensorType.Q4_0, GgufTensorType.Q4_0, GgufTensorType.Q4_0)) - .isTrue(); - } - } - - @Test - void acceptsAnExplicitFixedExecutionBatch() { - try (TornadoGgufBatchedMatrixKernel kernel = new TornadoGgufBatchedMatrixKernel(64)) { - assertThat(kernel.executionBatchSize()).isEqualTo(64); - assertThat(kernel.isEligible(GgufTensorType.Q4_0, 64, 3072, 1024)).isTrue(); - assertThat(kernel.isEligible(GgufTensorType.Q4_0, 65, 3072, 1024)).isFalse(); - } - } - - @Test - void decodeAccelerationIsExplicitAndUsesItsOwnSingleTokenShape() { - try (TornadoGgufBatchedMatrixKernel kernel = new TornadoGgufBatchedMatrixKernel(32, true)) { - assertThat(kernel.acceleratesDecode()).isTrue(); - assertThat(kernel.executionBatchSizeFor(1)).isEqualTo(1); - assertThat(kernel.executionBatchSizeFor(4)).isEqualTo(32); - assertThat(kernel.isEligible(GgufTensorType.Q4_0, 1, 3072, 1024)).isTrue(); - assertThat( - kernel.isDualEligible(GgufTensorType.Q4_0, 1024, GgufTensorType.Q4_0, 1024, 1, 1024)) - .isTrue(); - } - } - - @Test - void rejectsAnExecutionBatchBelowTheGpuThreshold() { - assertThatThrownBy(() -> new TornadoGgufBatchedMatrixKernel(3)) - .isInstanceOf(IllegalArgumentException.class) - .hasMessageContaining("executionBatchSize"); - } - - @Test - void padsTheLastPromptChunkWithoutReusingStaleActivations() { - float[] padded = {9.0f, 9.0f, 9.0f, 9.0f, 9.0f, 9.0f, 9.0f, 9.0f}; - - float[] executionInput = - TornadoGgufBatchedMatrixKernel.prepareExecutionInput( - new float[] {1.0f, 2.0f, 3.0f, 4.0f}, padded, 2, 4, 2); - - assertThat(executionInput).isSameAs(padded); - assertThat(executionInput).containsExactly(1.0f, 2.0f, 3.0f, 4.0f, 0.0f, 0.0f, 0.0f, 0.0f); - } - - @Test - void recommendsTheGroupedProjectionPlanThatTheExperimentImplements() { - try (TornadoGgufBatchedMatrixKernel kernel = new TornadoGgufBatchedMatrixKernel()) { - assertThat(kernel.planRecommendations()) - .containsEntry(PureJavaPlanConfiguration.GROUPED_PROJECTIONS_PROPERTY, "true") - .containsEntry(PureJavaPlanConfiguration.STAGED_QUANTIZED_FFN_PROPERTY, "false") - .containsEntry(PureJavaPlanConfiguration.STAGED_QUANTIZED_LAYER_PROPERTY, "false"); - } - } - - @Test - void admitsTheMixedKProjectionGroupThatQ4KMModelsPresent() { - try (TornadoGgufBatchedMatrixKernel kernel = new TornadoGgufBatchedMatrixKernel()) { - assertThat(kernel.planRecommendations()) - .containsEntry(PureJavaPlanConfiguration.MIXED_K_PROJECTIONS_PROPERTY, "true"); - assertThat( - kernel.supportsTriple(GgufTensorType.Q4_K, GgufTensorType.Q4_K, GgufTensorType.Q6_K)) - .isTrue(); - assertThat( - kernel.isTripleEligible( - GgufTensorType.Q4_K, - 2048, - GgufTensorType.Q4_K, - 512, - GgufTensorType.Q6_K, - 512, - 8, - 2048)) - .isTrue(); - } - } - - @Test - void keepsTheTwoActivationFamiliesOutOfOneGroupedDispatch() { - try (TornadoGgufBatchedMatrixKernel kernel = new TornadoGgufBatchedMatrixKernel()) { - // Q4_0 needs Q8_0 activations and the K-quants need Q8_K, so one prepared activation can - // never serve both. Mixed groups must fall back rather than silently use the wrong scales. - assertThat(kernel.supportsDual(GgufTensorType.Q4_0, GgufTensorType.Q4_K)).isFalse(); - assertThat(kernel.supportsDual(GgufTensorType.Q4_K, GgufTensorType.Q4_0)).isFalse(); - assertThat( - kernel.supportsTriple(GgufTensorType.Q4_K, GgufTensorType.Q4_0, GgufTensorType.Q4_K)) - .isFalse(); - // Q6_K has no dual kernel and is only ever the third matrix of a grouped dispatch. - assertThat(kernel.supportsDual(GgufTensorType.Q6_K, GgufTensorType.Q6_K)).isFalse(); - assertThat( - kernel.supportsTriple(GgufTensorType.Q6_K, GgufTensorType.Q4_K, GgufTensorType.Q4_K)) - .isFalse(); - assertThat(kernel.supportsDual(GgufTensorType.Q4_K, GgufTensorType.Q4_K)).isTrue(); - assertThat(kernel.supportsDual(GgufTensorType.Q5_K, GgufTensorType.Q5_K)).isFalse(); - } - } - - @Test - void requiresWholeSuperBlocksForKQuantProjections() { - try (TornadoGgufBatchedMatrixKernel kernel = new TornadoGgufBatchedMatrixKernel()) { - assertThat(kernel.isEligible(GgufTensorType.Q4_K, 8, 4096, 1024)).isTrue(); - assertThat(kernel.isEligible(GgufTensorType.Q6_K, 8, 4096, 1024)).isTrue(); - // 1120 is a multiple of 32 but not of 256: legal for Q4_0, never for a K-quant. - assertThat(kernel.isEligible(GgufTensorType.Q4_0, 8, 4096, 1120)).isTrue(); - assertThat(kernel.isEligible(GgufTensorType.Q4_K, 8, 4096, 1120)).isFalse(); - assertThat( - kernel.isDualEligible(GgufTensorType.Q4_K, 4096, GgufTensorType.Q4_K, 4096, 8, 1120)) - .isFalse(); - } - } - - @Test - void refusesTensorsTooLargeForAnIntIndexedDeviceBuffer() { - try (TornadoGgufBatchedMatrixKernel kernel = new TornadoGgufBatchedMatrixKernel()) { - // Q4_K stores 144 bytes per 256 values, so a tensor needs about 3.8e9 values before it - // stops fitting an int-indexed device buffer. A 27B-class vocabulary projection - // (262144 x 5120 = 755 MiB of Q4_K) is comfortably inside that; a tensor four times - // taller is not, and must stay on the Vector API rather than fail the load. - assertThat(kernel.isEligible(GgufTensorType.Q4_K, 8, 262_144, 5120)).isTrue(); - assertThat(kernel.isEligible(GgufTensorType.Q4_K, 8, 1_000_000, 5120)).isFalse(); - assertThat(kernel.isEligible(GgufTensorType.Q6_K, 8, 1_000_000, 5120)).isFalse(); - assertThat(kernel.isEligible(GgufTensorType.Q4_0, 8, 4_000_000, 5120)).isFalse(); - assertThat( - kernel.isTripleEligible( - GgufTensorType.Q4_K, - 1_000_000, - GgufTensorType.Q4_K, - 512, - GgufTensorType.Q6_K, - 512, - 8, - 5120)) - .isFalse(); - } - } - - @Test - void reportsNoRoutedProjectionsBeforeAnyDispatch() { - try (TornadoGgufBatchedMatrixKernel kernel = new TornadoGgufBatchedMatrixKernel()) { - assertThat(kernel.routedProjectionsByFormat()).isEmpty(); - assertThat(kernel.projectionPlanCount()).isZero(); - assertThat(kernel.calls()).isZero(); - assertThat(kernel.totalMillis()).isZero(); - } - } - - @Test - void refusesToDispatchProjectionsItDeclaredIneligible() { - try (TornadoGgufBatchedMatrixKernel kernel = new TornadoGgufBatchedMatrixKernel()) { - assertThatThrownBy( - () -> - kernel.multiply( - new float[8], - new float[8], - MemorySegment.ofArray(new byte[8]), - GgufTensorType.Q5_K, - 8, - 1, - 1)) - .isInstanceOf(UnsupportedOperationException.class) - .hasMessageContaining("not eligible"); - } - } -} diff --git a/benchmark-results/2026-10-07-g1-parity/README.md b/benchmark-results/2026-10-07-g1-parity/README.md new file mode 100644 index 000000000..83696ba1c --- /dev/null +++ b/benchmark-results/2026-10-07-g1-parity/README.md @@ -0,0 +1,60 @@ +# G1 passes: 1280 token ids identical, device against CPU + +**Measured 2026-10-07 on an RTX 4090 (compute capability 8.9, driver 13020), Runpod.** Models +revision `15661fb2836967e200322e5b0ceaa91416790994`. Kernel sha256 `4c534e622452df95...`, PTX target +`sm_80`. Model: Granite 4.1 3B Q4_K_M, sha256 recomputed by the gate from the file it opened. + +``` +PASS cuda-kernel-gate mode=parity device=NVIDIA GeForce RTX 4090 prompts=20 tokens=1280 identical +``` + +| gate | result | +| --- | --- | +| `tokenParity` (G1) | **passed** -- 1280 token ids identical across 20 prompts | +| `routingObservability` (G2) | passed -- 2,405,576 accelerated operations | +| `startupHonesty` (G5) | passed -- 344 ms readiness against a 120,000 ms ceiling | +| `qualified` | **true** | + +`parity.firstDivergence` is **null**: every token id sequence matched. `parity.selfTest` is +**false**, which is the field that makes this G1 evidence at all -- off-device both arms are the +Vector API and the command says `SELF-TEST` instead of `PASS`. One sequence hit end-of-generation, +which is recorded and deliberately not obeyed, so the compared length is not shortened where the +arms agree. + +`routing.totalDeclinedProjections` is **0**, from the counter added the same day. Nothing fell back +to the Java path, so the 1280 identical tokens were produced with the projections actually on the +device rather than by a quiet fallback that would have made parity trivial. + +## What changed since G1 last failed + +A note from a 2026-09-26 A40 run recorded G1 failing on an argmax flip at token 7, attributed to +`models_gqa_decode_attention`'s `expf`. Between then and now `expf` was rewritten on a different +principle, stated in `backend-cuda/src/main/rust/models-cuda-kernels/src/attention.rs`: + +> *"The requirement is not that it be accurate -- it is that it be **the same function the CPU +> runs**."* + +It transcribes the CPU's clamp, magic-constant rounding, Cody-Waite split, Taylor coefficients and +single-step exponent construction, in that order, with `fma` exactly where the CPU has one. That is +what exact token parity needed, and accuracy alone would not have given it. + +Two differences from the failing run are recorded rather than waved at: this is an RTX 4090 at +compute capability 8.9 where that was an A40 at 8.6, and this revision routes the FFN gate and up +projections to the device where that one left them on the CPU. The same `sm_80` module serves both +capabilities, but a single host does not establish hardware independence. **Re-run on an A40 or +L40S before claiming G1 holds across the qualifying profiles.** + +## Both gates now pass + +- **G1**, here: 1280/1280 token ids identical. +- **G4**, `benchmark-results/2026-10-07-g4-dualpath`: 6.184x decode against a 3.00x gate, 31.26 + against 5.06 tok/s, with a control arm that moved 0.2%. + +Numerically exact and 6.18x, on one host, on one model. + +## What this does not establish + +One model and one device. Granite 4.1 3B is 40 blocks of `granite` architecture with Q4_K and Q6_K +projections; it exercises neither a mixture-of-experts routed FFN nor a per-layer feed-forward width +nor an architecture whose attention scale differs. G1 is a no-tolerance gate, so each qualifying +hardware profile and each architecture family earns its own run. diff --git a/benchmark-results/2026-10-07-g1-parity/rust-ptx-parity.json b/benchmark-results/2026-10-07-g1-parity/rust-ptx-parity.json new file mode 100644 index 000000000..16259846a --- /dev/null +++ b/benchmark-results/2026-10-07-g1-parity/rust-ptx-parity.json @@ -0,0 +1,119 @@ +{ + "schemaVersion" : 2, + "createdAt" : "2026-10-07T19:29:53.432864217Z", + "policyVersion" : "gpu-large-model-rust-ptx-v2", + "modelsRevision" : "15661fb2836967e200322e5b0ceaa91416790994", + "mode" : "parity", + "accelerated" : true, + "refusalReason" : "", + "device" : { + "name" : "NVIDIA GeForce RTX 4090", + "computeCapability" : 89, + "globalMemoryBytes" : 25252724736, + "driverVersion" : 13020 + }, + "kernel" : { + "sha256" : "4c534e622452df95ca7fbb3accc65246be54e2c74d53714d98a03be6d31e8e0e", + "ptxTarget" : "sm_80", + "toolchain" : "nightly-2026-09-17", + "entryPoints" : [ "models_q4k_decode_projection", "models_q6k_decode_projection", "models_gqa_decode_attention" ] + }, + "environment" : { + "host" : "9342571901f5", + "osName" : "Linux", + "osVersion" : "6.8.0-146-generic", + "architecture" : "amd64", + "cpuModel" : "AMD EPYC 7452 32-Core Processor", + "processors" : 16, + "physicalMemoryBytes" : 142999998464, + "maxHeapBytes" : 32178700288, + "javaVersion" : "25.0.4.1", + "javaVendor" : "Eclipse Adoptium", + "vmName" : "OpenJDK 64-Bit Server VM" + }, + "configuration" : { + "mode" : "parity", + "modelPath" : "/work/model.gguf", + "modelSha256" : "662b0626cd58f443baea23559b469df6576a81d349649c59413b36a9fb32eb29", + "modelBytes" : 2099501664, + "promptsPath" : "/work/models/benchmark-results/2026-09-18-gpu-large-model/prompts.txt", + "promptsSha256" : "d0c0267f1eeb96b97a6e26f964d4e06c0e967ad7567e40881e54c791d74c7016", + "promptCount" : 20, + "maxTokens" : 64, + "warmupTokens" : 16, + "contextLength" : 4096, + "arm" : "both", + "sampling" : "greedy-argmax-lowest-index-v1", + "cudaDisabledProperty" : false, + "cudaDeviceOrdinal" : 0, + "jvmArguments" : [ "--add-modules=jdk.incubator.vector", "--enable-native-access=ALL-UNNAMED", "-XX:NativeMemoryTracking=summary" ], + "workingDirectory" : "/work/models" + }, + "routing" : { + "acceleratedOperations" : { + "F32/DECODE_ATTENTION" : 2040000, + "Q6_K/DECODE_PROJECTION" : 52296, + "Q4_K/PREFILL_PROJECTION" : 6240, + "Q6_K/PREFILL_PROJECTION" : 1040, + "Q4_K/DECODE_PROJECTION" : 306000 + }, + "refusals" : { }, + "declinedProjections" : { }, + "totalAcceleratedOperations" : 2405576, + "totalDeclinedProjections" : 0, + "inert" : false, + "kernelLaunches" : 416576, + "hostToDeviceTransfers" : 260456, + "deviceToHostTransfers" : 416576, + "hostToDeviceBytes" : 13756064960, + "deviceToHostBytes" : 8247738368, + "weightUploads" : 281, + "weightUploadBytes" : 2095104000, + "decodeSteps" : 1260, + "decodeProjections" : 358296, + "measuredDecodeSteps" : 1260, + "decodeStepsMarked" : true, + "launchesPerDecodeStep" : 321.0, + "transfersPerDecodeStep" : 522.0, + "activationBytesPerDecodeStep" : 1.5398952E7 + }, + "readinessMillis" : 344, + "parity" : { + "selfTest" : false, + "promptCount" : 20, + "tokensPerPrompt" : 64, + "comparedTokens" : 1280, + "sequencesHittingEndOfGeneration" : 1, + "identical" : true, + "firstDivergence" : null, + "promptDigests" : [ "bfe7bee92839632f1d68ebe8b65fb9cd3824ba3c926f8c6b0ccebec3571c985e", "9c98531297412ba6e3ebd3fc1bb3881552611d775b01b386d5c52258d96d4527", "decc00c635aa1608bc8affe58c034ea40ad5afea8712fcdcf4b350e1083d7b25", "6fae40b90f5a86c553ea914279f5c38de0fd8c6928d6e58b586c9af6e9e1c68d", "c6c78dc688a4443e094949c420f7e5e363ce737e52abf84e66f2aa2a20b83bc1", "a65c124623f65e639d05a06e5056d532e29532ae1952bdfeb408706f566834ab", "5f97c648fc5329865dac35d5fb20bb15667cd5a500c7752d60241db191aa7646", "3c61cb0871e38014fd1ffeab1e7c9b86fecae87dd39f8cc51daaeaf0ce3d5b1d", "85fa275abad2fcf7652db2060a0ca5ae85f2cb3ee895dbf9fdad3844277f215e", "c3d67901cd73e8dc1f6477e2da2316d0ceeb5c18d7a1ca2817c5e353936d9e20", "e3fb263583e2d93a9536dcd9bb735bb68e3d71c4e14ff704a7a5cea7abdb95ab", "55a4be086aef659ef78f2cf8d194c7aff228c95a2b23241a166b9f3842cb29e1", "0ba7006470ec8d2c1ec60056471ac68f000c8a3597c925b4497f76cc33c494f1", "531fe0c1a690b066dacf2ed30e63cd86737cb06615bb99a39f5d302d7f65d5e7", "16422f3a476e41cbe64164b421b9b66f97d07e82a22c240cff332ba328b81bd0", "00a434f63a5b3b89005ec8e98e7eb37ed2dc5fa6300533165b31df949314b102", "574cd13af5ceb84475eab456207f3cb0b258760a79a4228c631dabe5c0022388", "72a215d0a7efd7fae7dadc7b633701204d9f4bafe4487f302f5c13877aa9a444", "9fd82302b33caf3033c421ef429e69b273f487fac4f4ba2744905e26511c3a24", "9618c19f55a9736ee7cc5cac3f521d92762e3c1a28226e83194008b9b0c27ee4" ] + }, + "decode" : null, + "memory" : { + "peakProcessRssBytes" : 3707437056, + "residentBytes" : 1690697728, + "heapUsedBytes" : 377563128, + "peakDeviceBytes" : 2097012576, + "deviceTotalMemoryBytes" : 25252724736, + "reading" : "peakDeviceBytes covers kernel-owned allocations only (weights, staging scratch, per-call attention buffers); the CUDA context and loaded module are not included, so it is a lower bound. peakProcessRssBytes is process-wide and therefore not separable per arm in a two-arm run: use --arm accelerated and --arm control in separate processes for per-arm host peaks." + }, + "gates" : { + "routingObservability" : { + "gate" : "G2", + "passed" : true, + "evidence" : "2405576 accelerated operations" + }, + "startupHonesty" : { + "gate" : "G5", + "passed" : true, + "evidence" : "344 ms readiness against a 120000 ms ceiling" + }, + "tokenParity" : { + "gate" : "G1", + "passed" : true, + "evidence" : "1280 token ids identical across 20 prompts" + }, + "decodeSpeedup" : null + }, + "qualified" : true +} diff --git a/benchmark-results/2026-10-07-g1-parity/worker.log b/benchmark-results/2026-10-07-g1-parity/worker.log new file mode 100644 index 000000000..b8471508a --- /dev/null +++ b/benchmark-results/2026-10-07-g1-parity/worker.log @@ -0,0 +1,71 @@ +[19:21:04] nvidia-smi: +NVIDIA GeForce RTX 4090, 8.9, 595.91.07 +[19:21:04] installing jdk 25 +[19:21:09] jdk: openjdk version "25.0.4.1" 2026-08-18 LTS +[19:21:09] installing pinned rust nightly-2026-09-17 +[19:21:30] rust: rustc 1.100.0-nightly (923c95cdf 2026-09-16) +[19:21:30] fetching source +[19:21:33] source at /work/models, revision 15661fb2836967e200322e5b0ceaa91416790994 +[19:21:33] === STEP 1: ptxas assembly (cheap gate) === + Downloaded getopts v0.2.24 + Downloaded rustc-demangle v0.1.28 + Downloaded hashbrown v0.17.1 + Downloaded libc v0.2.189 + Compiling compiler_builtins v0.1.160 (/root/.rustup/toolchains/nightly-2026-09-17-x86_64-unknown-linux-gnu/lib/rustlib/src/rust/library/compiler-builtins/compiler-builtins) + Compiling core v0.0.0 (/root/.rustup/toolchains/nightly-2026-09-17-x86_64-unknown-linux-gnu/lib/rustlib/src/rust/library/core) + Compiling models-cuda-kernels v0.3.46 (/work/models/backend-cuda/src/main/rust/models-cuda-kernels) + Finished `release` profile [optimized] target(s) in 21.85s +[19:23:48] ptx: backend-cuda/build/generated/cuda-resources/META-INF/models/cuda/models-cuda-kernels.ptx (64104 bytes) +[19:23:48] ptxas OK +[19:23:48] === STEP 2: capability gate === +WARNING: Using incubator modules: jdk.incubator.vector +PASS cuda-kernel-gate mode=capability device=NVIDIA GeForce RTX 4090 cc=89 kernel=4c534e622452 readiness=527 ms report=/work/capability.json +[19:23:50] capability report uploaded +[19:23:50] === STEP 2a: Q6_K device parity (gate before any model) === +CudaQ6KDeviceParityTest > a Q6_K row projection is bit-exact against the CPU control at every width > 1 super-block(s) PASSED + +CudaQ6KDeviceParityTest > a Q6_K row projection is bit-exact against the CPU control at every width > 2 super-block(s) PASSED + +CudaQ6KDeviceParityTest > a Q6_K row projection is bit-exact against the CPU control at every width > 3 super-block(s) PASSED + +CudaQ6KDeviceParityTest > a Q6_K row projection is bit-exact against the CPU control at every width > 32 super-block(s) PASSED + +CudaQ6KDeviceParityTest > a Q6_K row projection is bit-exact against the CPU control at every width > 33 super-block(s) PASSED + +CudaQ6KDeviceParityTest > a Q6_K row projection is bit-exact against the CPU control at every width > 48 super-block(s) PASSED + +BUILD SUCCESSFUL in 11s +13 actionable tasks: 3 executed, 10 up-to-date +Consider enabling configuration cache to speed up this build: https://docs.gradle.org/9.4.1/userguide/configuration_cache_enabling.html +[19:24:01] Q6_K device parity OK +[19:24:01] === fetching model (granite 4.1 3b q4_k_m, 1.96GB) === +[19:24:27] model sha verified +[19:24:27] === STEP 4: parity gate G1 === +WARNING: Using incubator modules: jdk.incubator.vector +Oct 07, 2026 7:24:30 PM com.integrallis.vectors.core.PanamaVectorUtilSupportProvider create +INFO: vectors-core: Using Panama Vector API SIMD provider (vector bits: 256) +Oct 07, 2026 7:24:30 PM com.integrallis.vectors.core.VectorizationProvider logActiveToggles +INFO: vectors-core: provider=PanamaVectorUtilSupport panama=true maxBits=256 preferredBits=256 fastVectorFMA=true fastScalarFMA=true sve=false ggufParallel=true ggufParallelThreshold=1048576 ggufExecutor=persistent ggufThreads=16 ggufChunksPerThread=2 mappedKQuantLongOffsets=auto(q4=true,q5=false,q6=false) q4ShortPairwiseSupported=true q4UnsignedPairwiseSupported=true toggles=[(defaults ? no -Dvectors.* overrides)] +PASS cuda-kernel-gate mode=parity device=NVIDIA GeForce RTX 4090 prompts=20 tokens=1280 identical report=/work/parity.json +[19:29:54] decode report uploaded +[19:29:54] parity gate rc=0 +[19:29:54] === the numbers this run exists for === + --- G1 verdict --- + accelerated = True + refusalReason = + selfTest = False + qualified = True + tokenParity: passed=True | 1280 token ids identical across 20 prompts + --- sequences --- + parity.promptCount = 20 + parity.sequencesHittingEndOfGeneration = 1 + --- first divergence --- + none: every token id sequence identical + --- routing --- + routing.kernelLaunches = 416576 + routing.decodeProjections = 358296 + routing.totalDeclinedProjections = 0 + routing.launchesPerDecodeStep = 321.0 + routing.transfersPerDecodeStep = 522.0 + declined: none +[19:29:54] FINAL_STATUS=COMPLETE diff --git a/benchmark-results/2026-10-07-tornado-removal/README.md b/benchmark-results/2026-10-07-tornado-removal/README.md new file mode 100644 index 000000000..d7c2a9180 --- /dev/null +++ b/benchmark-results/2026-10-07-tornado-removal/README.md @@ -0,0 +1,106 @@ +# Why backend-tornado was removed + +**Decision 2026-10-07: Models ships one GPU implementation.** The measurement below is why the +TornadoVM arm is not it. Same RTX 4090, same model, same day as `backend-cuda`'s gates: + +| | `backend-cuda` | `backend-tornado` | +| --- | --- | --- | +| G1 exact token parity | **1280 / 1280 identical** | not reached | +| decode | **31.26 tok/s** | 6.49 tok/s | +| CPU control, same host | 5.06 tok/s | 5.06 tok/s | +| readiness | 344 ms | 32,171 ms | +| device launch errors | 0 | **36 x CUDA 701** | +| self-contained artifact | yes | no -- needs a separately installed runtime | + +A second implementation of the same capability, five times slower than the first and barely ahead of +the CPU path it is meant to accelerate, is not worth the surface it costs. It is removed rather than +carried: `backend-tornado`, its publication-allowlist entry, the `accelerator-profile` command it +drove, its TornadoVM benchmark arm, and the release precondition that pointed at its results +directory. + +The original run notes follow, kept because the numbers above come from them and because the +`cuLaunchKernel` failures are the substantive finding. + +## The run as it was recorded + +**Measured 2026-10-07.** RTX 4090, compute capability 8.9. TornadoVM **v5.2.0-jdk25** built from +source with the PTX backend, JDK 25. Models revision `10af16c33be7e023c114812e6137b47fab7e0c90`. +Model: Granite 4.1 3B Q4_K_M, sha256 verified on download. The `RELEASING.md` precondition names +this directory, so the run is retained here either way. + +Command, exactly as `models-bench/README.md` documents it: + +``` +tornado -cp "$classpath" \ + --params="accelerator-profile --model /work/model.gguf --tokens 64 --batch 32 --require true \ + --output /work/accelerator-profile.json" \ + com.integrallis.models.bench.InferenceBenchmarkCli +``` + +## What it reported + +``` +accelerator profile: accelerated=true device=cuda-0 reason=eligible readiness=32171 ms plans=321 + prompt=26 tokens prefill=69.34 tok/s (375.0 ms) decode=64 tokens 6.49 tok/s (9857.8 ms) + routed projections by format: {Q4_K=16800, Q6_K=2870} +``` + +Exit status 0. `--require true` was set, so an ineligible accelerator would have been a hard +failure; it was not reached. + +**Both format counts are non-zero**, which is the check the README asks for: *"a run whose `Q6_K` +count is zero routed only part of the model while still producing a plausible throughput number."* +Q4_K 16,800 and Q6_K 2,870, so the mixed-format trap is not what happened here. + +## What it does not report, and why this is not a clean pass + +The run emitted **36 occurrences** of + +``` +[TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 +``` + +CUDA 701 is `CUDA_ERROR_LAUNCH_OUT_OF_RESOURCES`. Thirty-six kernel launches failed on the device. + +**The report JSON records none of them.** Its keys are `accelerated`, `decodeMillis`, +`decodeTokensPerSecond`, `device`, `generatedTokens`, `model`, `prefillMillis`, +`prefillTokensPerSecond`, `projectionPlans`, `promptTokens`, `readinessMillis`, `reason`, +`requiredBytes`, `routedProjectionsByFormat`, `timestamp`. There is no field for a failed launch, so +`accelerated: true` and `decodeTokensPerSecond: 6.49` are what a reader of the artifact alone would +take away. + +So the gate's exit status is not evidence that the device path worked. **Treat this as a failed +precondition with a known cause, not as a pass**, and do not file the throughput figure as a +measurement of accelerated decode. + +The throughput supports that reading rather than contradicting it. Decode at **6.49 tok/s** is +barely above the **5.06 tok/s** Vector API control measured on the same model and the same GPU class +in `benchmark-results/2026-10-07-g4-dualpath`, where the `backend-cuda` arm reached **31.26 tok/s**. +A device arm that errors 36 times and lands within 30% of its own CPU control is behaving like one +whose launches mostly failed. + +Readiness was **32,171 ms** against G5's 120,000 ms ceiling -- inside budget, and two orders of +magnitude above `backend-cuda`'s 344-353 ms on the same hardware. + +## The pattern, which is the transferable part + +This is the third instance in one day of evidence reporting success while something silently did not +run: + +1. `backend-cuda` left the FFN gate and up projections on the CPU -- 53.3% of a layer's projection + arithmetic -- because `multiplyDual` was unimplemented, and an empty `refusals` map read as + "nothing fell back". +2. The counter added to catch that printed `None`, because it was never serialised into the report. +3. Here, 36 failed `cuLaunchKernel` calls with `accelerated: true` and no field to hold them. + +The fix in each case is the same shape: the artifact has to carry the negative. `accelerator-profile` +needs a launch-failure count, and `--require true` should fail on a nonzero one. + +## Next, in order + +1. Give `accelerator-profile` a launch-failure counter and make `--require true` respect it, so this + cannot read as a pass again. +2. Diagnose the 701 itself. `plans=321` retained execution plans and a 32-second readiness suggest + per-plan device resources rather than model size; `--batch 32` and the retained-plan count are the + first two things to vary. +3. Re-run only after both, and only then record a `backend-tornado` verdict. diff --git a/benchmark-results/2026-10-07-tornado-removal/accelerator-profile.json b/benchmark-results/2026-10-07-tornado-removal/accelerator-profile.json new file mode 100644 index 000000000..4743db3a2 --- /dev/null +++ b/benchmark-results/2026-10-07-tornado-removal/accelerator-profile.json @@ -0,0 +1,20 @@ +{ + "timestamp" : "2026-10-07T20:33:51.220769153Z", + "model" : "/work/model.gguf", + "accelerated" : true, + "device" : "cuda-0", + "reason" : "eligible", + "requiredBytes" : 4467438784, + "readinessMillis" : 32171, + "projectionPlans" : 321, + "routedProjectionsByFormat" : { + "Q4_K" : 16800, + "Q6_K" : 2870 + }, + "promptTokens" : 26, + "generatedTokens" : 64, + "prefillTokensPerSecond" : 69.3391006823717, + "decodeTokensPerSecond" : 6.492315709531025, + "prefillMillis" : 374.968809, + "decodeMillis" : 9857.807732 +} \ No newline at end of file diff --git a/benchmark-results/2026-10-07-tornado-removal/worker.log b/benchmark-results/2026-10-07-tornado-removal/worker.log new file mode 100644 index 000000000..fc800b180 --- /dev/null +++ b/benchmark-results/2026-10-07-tornado-removal/worker.log @@ -0,0 +1,121 @@ +[20:22:39] nvidia-smi: +NVIDIA GeForce RTX 4090, 8.9, 570.158.01 +Cuda compilation tools, release 12.8, V12.8.93 +Build cuda_12.8.r12.8/compiler.35583870_0 +[20:22:39] installing build prerequisites +[20:26:06] maven: Apache Maven 3.8.7 +[20:26:06] cmake: cmake version 3.28.3 +[20:26:06] installing the pinned rust nightly -- :models-bench:installDist pulls in +[20:26:06] :backend-cuda:compilePtx, which needs cargo and -Zbuild-std's rust-src +[20:26:24] cargo: cargo 1.100.0-nightly (495c385d0 2026-09-16) +[20:26:24] installing jdk 25 +[20:26:28] jdk: openjdk version "25.0.4.1" 2026-08-18 LTS +[20:26:28] === building TornadoVM v5.2.0-jdk25 from source, PTX backend === +[20:26:28] TORNADOVM_HOME=/work/tornado-sdk (an input to bin/compile, not an output) +[20:26:30] tornado tree at v5.2.0-jdk25 +[20:26:30] building via bin/compile --jdk jdk25 --backend ptx (bin/compile builds -P,; the wrong jdk value gives -source 8) +[20:27:58] tornado build rc=0 +20:27:58 [INFO] tornado-cufft ...................................... SUCCESS [ 1.567 s] +20:27:58 [INFO] tornado-cudnn ...................................... SUCCESS [ 1.570 s] +20:27:58 [INFO] tornado-cusparse ................................... SUCCESS [ 1.151 s] +20:27:58 [INFO] tornado-cutlass .................................... SUCCESS [ 1.260 s] +20:27:58 [INFO] tornado-unittests .................................. SUCCESS [ 5.187 s] +20:27:58 [INFO] tornado-annotation ................................. SUCCESS [ 1.372 s] +20:27:58 [INFO] tornado-assembly ................................... SUCCESS [ 28.687 s] +20:27:58 [INFO] ------------------------------------------------------------------------ +20:27:58 [INFO] BUILD SUCCESS +20:27:58 [INFO] ------------------------------------------------------------------------ +20:27:58 [INFO] Total time: 01:05 min (Wall Clock) +20:27:58 [INFO] Finished at: 2026-10-07T20:27:58Z +20:27:58 [INFO] ------------------------------------------------------------------------ +Maven build succeeded +########################################################################### +TornadoVM build success +Updating PATH and TORNADOVM_HOME to tornadovm-5.2.0-jdk25-ptx-linux-amd64 +Backend : PTX +Commit : db146ae +########################################################################### +Generated utilities: + [Unix-env]: /work/tornadovm/setvars.sh + [argfile]: /work/tornadovm/dist/tornadovm-5.2.0-jdk25-ptx-linux-amd64/tornadovm-5.2.0-jdk25-ptx/tornado-argfile +[INFO] Run first: source setvars.sh + +[20:27:58] not at $TORNADOVM_HOME/bin/tornado, searching the assembled dist +[20:27:58] launcher=/work/tornadovm/dist/tornadovm-5.2.0-jdk25-ptx-linux-amd64/tornadovm-5.2.0-jdk25-ptx/bin/tornado +[20:27:58] TORNADOVM_HOME now /work/tornadovm/dist/tornadovm-5.2.0-jdk25-ptx-linux-amd64/tornadovm-5.2.0-jdk25-ptx +[20:27:58] backends: tornado.backends=ptx-backend +version=5.2.0-jdk25 +branch=UNKNOWN +commit=db146ae +WARNING: Using incubator modules: jdk.incubator.vector + +Number of Tornado drivers: 1 +Driver: PTX + Total number of PTX devices : 1 + Tornado device=0:0 (DEFAULT) + PTX -- PTX -- NVIDIA GeForce RTX 4090 + Global Memory Size: 23.5 GB + Local Memory Size: 48.0 KB + Workgroup Dimensions: 3 + Total Number of Block Threads: [2147483647, 65535, 65535] + Max WorkGroup Configuration: [1024, 1024, 64] + Device OpenCL C version: N/A + + +[20:27:59] === fetching models source === +[20:28:03] models at 10af16c33be7e023c114812e6137b47fab7e0c90 +[20:28:03] === building models-bench dist === +[20:29:54] installDist rc=0 +[20:29:54] === fetching model === +[20:33:06] model sha verified +[20:33:06] === accelerator-profile, the documented gate command === +[20:33:52] accelerator-profile rc=0 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 + [TornadoVM-PTX-JNI] ERROR : cuLaunchKernel -> Returned: 701 +accelerator profile: accelerated=true device=cuda-0 reason=eligible readiness=32171 ms plans=321 + prompt=26 tokens prefill=69.34 tok/s (375.0 ms) decode=64 tokens 6.49 tok/s (9857.8 ms) + routed projections by format: {Q4_K=16800, Q6_K=2870} +report: /work/accelerator-profile.json +[20:33:53] profile uploaded +[20:33:53] === per-format routing, the point of the command === + accelerated = True reason = eligible + device = cuda-0 + readiness = 32171 + prefillTokensPerSecond = 69.3391006823717 + decodeTokensPerSecond = 6.492315709531025 + routedProjectionsByFormat = {"Q4_K": 16800, "Q6_K": 2870} +[20:33:53] FINAL_STATUS=COMPLETE diff --git a/build.gradle.kts b/build.gradle.kts index e94122378..333781912 100644 --- a/build.gradle.kts +++ b/build.gradle.kts @@ -51,8 +51,8 @@ val publishedModuleNames = "models-rag", "models-semantic-order", "backend-java", - "backend-tornado", "backend-native", + "backend-cuda", "backend-apple", "models-langchain4j", "models-spring-ai", diff --git a/docs/content/antora.yml b/docs/content/antora.yml index 4dcabe8f6..b7787933d 100644 --- a/docs/content/antora.yml +++ b/docs/content/antora.yml @@ -1,7 +1,7 @@ name: models title: Models version: 'current' -display_version: '0.3.52' +display_version: '0.3.53' prerelease: false start_page: ROOT:index.adoc nav: @@ -10,7 +10,7 @@ asciidoc: attributes: source-language: java source-highlighter: highlight.js - models-version: '0.3.52' + models-version: '0.3.53' modeljars-version: '0.1.31' vectors-version: '0.1.28' url-models-github: https://github.com/integrallis/models diff --git a/docs/content/modules/ROOT/pages/architecture.adoc b/docs/content/modules/ROOT/pages/architecture.adoc index dbd5e113f..a05069c4b 100644 --- a/docs/content/modules/ROOT/pages/architecture.adoc +++ b/docs/content/modules/ROOT/pages/architecture.adoc @@ -30,10 +30,12 @@ loaded through Java's final FFM API. It does not embed or call llama.cpp or Ollama. The Java graph remains authoritative, so the backend boundary is narrow and testable. -`backend-tornado` is a Java-authored device path for that same graph. TornadoVM -compiles eligible Q4_0 projections from Java bytecode for a qualified NVIDIA -GPU; attention and unsupported projections remain on the Vector API. No model -server or second inference engine is involved. +`backend-cuda` is the device path for that same graph. Models-owned Rust kernels +compiled to PTX run the eligible K-quant projections and grouped-query decode +attention on a qualified NVIDIA GPU; unsupported formats and shapes remain on +the Vector API, and a declined projection is counted rather than silently +dropped. No model server or second inference engine is involved, and Models +ships only this one GPU implementation. `backend-apple` is also in-process: Java reaches Apple's on-device Foundation Models framework through a bundled, integrity-checked FFM diff --git a/docs/content/modules/ROOT/pages/gpu-acceleration.adoc b/docs/content/modules/ROOT/pages/gpu-acceleration.adoc index 53f0aea42..b570f4369 100644 --- a/docs/content/modules/ROOT/pages/gpu-acceleration.adoc +++ b/docs/content/modules/ROOT/pages/gpu-acceleration.adoc @@ -1,188 +1,107 @@ -= Java GPU Acceleration - -`backend-tornado` optionally executes Models-owned Java projection kernels -on a qualified NVIDIA GPU. The model parser, tokenizer, transformer graph, -attention, KV cache, sampling, and generation loop remain inside Models. There -is no external inference server and no Rust or handwritten CUDA kernel in this -path. - -== Qualified Scope - -The production selector currently admits: - -* GGUF Q4_0, Q4_K, and Q6_K projection work; -* NVIDIA devices exposed by TornadoVM's PTX backend; -* a device with enough memory for retained prefill and decode plans plus the - required safety margin; -* a single weight tensor small enough to address with a 32-bit index - (under 2 GiB on the device); and -* fixed 32-token prefill plans and separate single-token decode plans. - -Attention and unsupported tensors continue on the Vector API. AMD, Intel, and -Metal devices remain CPU fallback paths until each vendor path passes a real -hardware parity and performance gate. - -=== Weight Formats And Activation Families - -Q4_0 weights are multiplied by Q8_0 activations; Q4_K and Q6_K weights are -multiplied by Q8_K activations. A grouped dispatch shares one prepared -activation across two or three weight tensors, so the two families are never -combined inside one group even though a `Q4_K_M` model contains both. The -admitted grouped shapes are `Q4_0/Q4_0`, `Q4_0/Q4_0/Q4_0`, `Q4_K/Q4_K`, -`Q4_K/Q4_K/Q4_K`, and `Q4_K/Q4_K/Q6_K` — the last being the query, key, and -value group a `Q4_K_M` model presents, since llama.cpp promotes the value -projection to Q6_K. Any other combination falls back to the Vector API -per projection, and Q5_K, Q2_K, Q3_K, Q8_0, F32, and BF16 tensors are never -routed to the device. - -K-quant tensors additionally require a column count that is a whole multiple of -256, the K-quant super-block size. Q4_0's 32-value blocks are not sufficient. - -=== Numeric Parity With The CPU Kernels - -The K-quant kernels accumulate every per-super-block quantized dot product and -minimum correction in `int`, so those reductions are exact and independent of -how work is scheduled. Only the per-super-block scale application is floating -point, and it runs in ascending super-block order in the same two-step form the -vectors-core CPU kernels use. The CPU kernels fuse those steps with -`Math.fma` and the device kernels use a plain multiply and add, so the -guaranteed contract is agreement to within two float roundings per super-block. -Off-device parity tests assert both halves of this: bit-for-bit equality on -super-blocks whose scales are powers of two, and the rounding budget on -pseudo-random super-blocks. - -Device-side parity — that TornadoVM's PTX backend lowers these kernels to the -same arithmetic — is a hardware gate, not a unit test. - -== Large Models Are Refused, Not Guessed At - -The capacity gate adds up what a model would place on the device: weights under -the plan shape the kernel builds, per-plan scratch, and any device-resident KV -cache. It refuses a model when a term it needs is unknown rather than assuming -it is zero, so a model above 8 GiB of weights is not admitted on a file-size-only -budget. It also refuses a tensor at or above the 2 GiB TornadoVM `ByteArray` -limit, and a plan set whose eager readiness would exceed 120 seconds. - -A 27B-class Q4_K_M model does not run on this path today. It holds no Q4_0 -tensors, and the retained plan cache keeps one host and one device copy of every -weight per batch shape, so the shipped plan shape asks for roughly twice the -model on both sides. Ineligibility messages name what was needed, what was -available, and which plan shape would have fit, so a refusal is actionable -rather than a bare fallback. - -== Add The Optional Backend - -[source,kotlin,subs="attributes+"] ----- -dependencies { - implementation("com.integrallis:backend-tornado:{models-version}") -} ----- - -Install a TornadoVM distribution compatible with the application's JDK and -with its PTX backend enabled. Follow the -https://tornadovm.readthedocs.io/en/latest/installation.html[official TornadoVM installation guide^]. -The `backend-tornado` dependency provides the Models integration but does not -bundle or transitively install TornadoVM's device runtime. Device discovery -requires the TornadoVM distribution and launch configuration. Start the -application with `tornado`, or with ordinary `java` and TornadoVM's generated -argument file. += NVIDIA GPU Acceleration -When `backend-tornado` is on the application classpath, `PureJavaBackend.loadAutomatic(...)` -discovers it through Java's standard `ServiceLoader`. ModelJars uses that automatic -entry point for artifacts qualified for the Java backend, so adding the optional -dependency and launching with TornadoVM is sufficient; application inference code -does not change. Without the optional module or a qualified device, the same call -loads the Vector API backend. +`backend-cuda` optionally executes the K-quant projections and grouped-query +decode attention on a qualified NVIDIA GPU, using Models-owned Rust kernels +compiled to PTX. The model parser, tokenizer, transformer graph, KV cache, +sampling and generation loop remain inside Models. There is no external +inference server, no vendor math library and no handwritten CUDA C++. -== Open And Inspect The Backend +Models ships *one* GPU implementation. A TornadoVM-based arm existed until +0.3.53 and was removed: on the same RTX 4090 and the same model it reached +6.49 tok/s decode over 36 failed `cuLaunchKernel` calls, against 31.26 tok/s +for this path, and it required a separately installed device runtime that a +Maven artifact cannot carry. The measurement is retained in +`benchmark-results/2026-10-07-tornado-removal`. -[source,java] ----- -var options = TornadoBackendOptions.defaults(); +== Nothing activates it by accident -try (var runtime = TornadoBackend.open(modelPath, backendConfiguration, options)) { - var backend = runtime.backend(); - var status = runtime.status(); - - System.out.printf("device=%s accelerated=%s readiness=%s%n", - status.device(), status.accelerated(), status.readinessTime()); - // Pass backend to GenerationLoop or the normal Models runtime pipeline. -} ----- - -Applications that do not need the detailed readiness status can use the -service-loaded entry point directly: +The jar carries no `META-INF/services` entry, so no `ServiceLoader` discovers +it. An application opens the accelerator explicitly: [source,java] ---- -try (var backend = PureJavaBackend.loadAutomatic(modelPath, backendConfiguration)) { - // Use the normal Models generation pipeline. -} +CudaGgufBatchedMatrixKernel.Status status = CudaGgufBatchedMatrixKernel.open(); ---- -Defaults perform eager readiness, accelerate eligible prefill and decode -projections, and safely load `PureJavaBackend` when acceleration is not -available. A qualification or deployment gate can set `requireAccelerator` to -`true`; in that mode an ineligible device or initialization failure stops the -load instead of falling back. +`open()` never throws for an absent, old or ineligible device; it returns a +status explaining why. `-Dmodels.cuda.disabled=true` refuses unconditionally, +and `-Dmodels.cuda.attention.disabled=true` ablates only the attention kernel, +counting the refusal rather than hiding it. -Readiness is intentional. TornadoVM retains compiled code with each execution -plan, so Models constructs and compiles the fixed shapes before accepting a -visible request. `TornadoBackendStatus` reports that cost and the selected -device. +== One module for every qualifying device -A mixed-format model can look accelerated while one of its formats silently -falls back, so `TornadoBackendRuntime` also reports how many projections reached -the device per GGUF weight format. A grouped dispatch counts once per matrix. +PTX is device code, so a single `sm_80` module serves every device of compute +capability 8.0 or above. It travels inside the jar under +`META-INF/models/cuda/` with a SHA-256 the Java loader recomputes on load, and +an ABI field that must match the Java binding. No CUDA toolkit is needed to +build or to run; only a shipping `libcuda.so.1`. -[source,java] ----- -try (var runtime = TornadoBackend.open(modelPath, backendConfiguration, options)) { - // Run a prefill first; counts are empty until projections are dispatched. - System.out.println(runtime.routedProjectionsByFormat()); // {Q4_K=1344, Q6_K=64} - System.out.println(runtime.projectionPlanCount()); -} ----- +== Measured gates -== Accelerator Profile Gate - -`models-bench` runs prefill and decode for a GGUF file on the Tornado backend and -writes the device, readiness cost, throughput, and per-format routing counts as -JSON: +Run them with `models-bench cuda-kernel-gate`, which has three modes and one +report shape: [source,shell] ---- -./gradlew :models-bench:installDist +models-bench cuda-kernel-gate --mode capability --require-device true \ + --report capability.json --models-revision "$(git rev-parse HEAD)" -lib=models-bench/build/install/models-bench/lib -classpath=$(printf '%s:' "$lib"/*.jar) +models-bench cuda-kernel-gate --mode parity --model model.gguf \ + --prompts prompts.txt --max-tokens 64 --require-device true \ + --report parity.json --models-revision "$(git rev-parse HEAD)" -tornado -cp "$classpath" \ - --params="accelerator-profile --model /path/to/model-Q4_K_M.gguf \ - --tokens 64 --batch 32 --require true \ - --output build/reports/inference/accelerator-profile.json" \ - com.integrallis.models.bench.InferenceBenchmarkCli +models-bench cuda-kernel-gate --mode decode --model model.gguf \ + --prompts prompts.txt --max-tokens 64 --warmup-tokens 16 --arm both \ + --require-device true --report decode.json \ + --models-revision "$(git rev-parse HEAD)" ---- -`--require true` makes an unavailable or ineligible accelerator a hard failure -rather than a silent Vector API fallback. Without a qualified device the command -still runs and the report records `accelerated=false` with the reason. - -== Measured Gates +`--arm both` measures the accelerated arm and the Vector API control in the +same process on the same host, which is the only denominator that counts. +`--models-revision` requires a full 40-character SHA. -The release candidate preserved the exact greedy CPU token sequence on both -qualified profiles using Qwen3 0.6B Q4_0. These are Q4_0 numbers; the K-quant -kernels have not yet been measured on hardware, and no throughput claim is made -for them here. +Measured on an RTX 4090 (compute capability 8.9), Granite 4.1 3B Q4_K_M, 20 +prompts: -[cols="1,1,1,1",options="header"] +[cols="2,1,3"] |=== -|Profile |Warm prefill |Median decode |Eager readiness -|NVIDIA A16-2Q |4.938 s (`1.93x` CPU) |73.991 ms (`1.21x` CPU) |14.278 s -|NVIDIA A40-4Q |2.150 s (`4.40x` CPU) |65.767 ms (`1.32x` CPU) |13.554 s +|Gate |Result |Evidence + +|G1 exact token parity +|*passed* +|1280 of 1280 token ids identical, `benchmark-results/2026-10-07-g1-parity` + +|G4 decode speed +|*passed* +|6.184x against a 3.00x threshold, 31.26 against 5.06 tok/s, + `benchmark-results/2026-10-07-g4-dualpath` + +|G2 routing observability +|passed +|2,405,576 accelerated operations, zero declined projections + +|G5 startup honesty +|passed +|344 ms readiness against a 120,000 ms ceiling |=== -These are qualification measurements for the tested virtual-GPU profiles, not -universal performance promises. Application latency depends on device shape, -CPU allocation, prompt length, model size, and driver/runtime versions. +G1 carries no tolerance. Each qualifying hardware profile and each +architecture family earns its own run before the result generalises: the figures +above are one device and one model. + +== What routes, and how to tell + +The report records accelerated operations per weight format and stage, plus +`declinedProjections` keyed by format, stage, reason and shape. That second map +exists because an unimplemented path is not a refusal: the FFN gate and up +projections ran on the CPU through an entire measurement because +`multiplyDual` was missing, 53.3% of a layer's projection arithmetic, and an +empty refusals map read as "everything ran on the device". A projection that +answers no to `isEligible` now increments a counter that names it. + +Read `launchesPerDecodeStep`, `transfersPerDecodeStep` and +`activationBytesPerDecodeStep` together with the throughput. The decode path +issues roughly 321 launches and 522 transfers per token on a 40-block model, +because every projection result returns to a Java array; that is the standing +overhead and the next structural lever against it is device-resident +activations. diff --git a/docs/content/modules/ROOT/pages/index.adoc b/docs/content/modules/ROOT/pages/index.adoc index 36f76c4ff..12f5e995f 100644 --- a/docs/content/modules/ROOT/pages/index.adoc +++ b/docs/content/modules/ROOT/pages/index.adoc @@ -27,10 +27,11 @@ the same Java pipeline but replaces selected, measured bottleneck kernels with a small Models-owned Rust library through Java's Foreign Function and Memory (FFM) API. It does not embed or invoke llama.cpp or Ollama. -The optional `backend-tornado` module keeps that graph and its projection -kernels in Java while TornadoVM compiles qualified Q4_0 work for an NVIDIA GPU. -It performs an explicit readiness pass and falls back to the Vector API when -the device, memory capacity, or artifact is not eligible. +The optional `backend-cuda` module keeps that graph in Java while Models-owned +Rust kernels, compiled to PTX and carried inside the jar, run the K-quant +projections and grouped-query decode attention on an NVIDIA GPU. It performs an +explicit readiness pass and falls back to the Vector API when the device, memory +capacity, or artifact is not eligible. New to local model execution? Start with xref:concepts.adoc[Inference Concepts]. @@ -104,8 +105,9 @@ image::models-0001.png[Models runtime architecture,align=center] |`backend-java` |GGUF, supported Safetensors, and CACT model loading with Java 25 Vector API execution -|`backend-tornado` -|Optional Java-authored Q4_0 projection acceleration on qualified NVIDIA GPUs +|`backend-cuda` +|Optional Rust-authored PTX projection and decode-attention acceleration on +qualified NVIDIA GPUs |`backend-native` |The Java backend with selective Rust/FFM bottleneck-kernel substitution diff --git a/docs/content/modules/ROOT/pages/modules.adoc b/docs/content/modules/ROOT/pages/modules.adoc index 331b1433a..b548e27d0 100644 --- a/docs/content/modules/ROOT/pages/modules.adoc +++ b/docs/content/modules/ROOT/pages/modules.adoc @@ -25,8 +25,9 @@ execution across local and hosted clients |`backend-java` |GGUF parser, tokenizers, model graph, KV cache, and Java kernels -|`backend-tornado` |Optional Java-authored Q4_0 projection kernels compiled -for qualified NVIDIA GPUs by TornadoVM, with eager readiness and CPU fallback +|`backend-cuda` |Optional Models-owned Rust kernels compiled to PTX for +qualified NVIDIA GPUs: K-quant projections and grouped-query decode attention, +opt-in, with CPU fallback |`backend-native` |Java 25 and Vector API backend with selective Rust kernel acceleration through FFM @@ -77,7 +78,7 @@ models-api <- models-runtime <- models <- backend-java <- vectors-core - <- backend-tornado <- TornadoVM device compiler/runtime + <- backend-cuda <- Models Rust PTX kernels through FFM <- backend-native <- Models Rust kernels through FFM <- backend-apple <- Apple's Foundation Models framework through FFM <- models-router <- vectors-db @@ -105,7 +106,7 @@ Application dependencies use: * `com.integrallis:models` * `com.integrallis:backend-java` -* `com.integrallis:backend-tornado` +* `com.integrallis:backend-cuda` * `com.integrallis:backend-native` * `com.integrallis:backend-apple` * `com.integrallis:models-audio` diff --git a/docs/landing/index.html b/docs/landing/index.html index b25aed447..933acd235 100644 --- a/docs/landing/index.html +++ b/docs/landing/index.html @@ -36,7 +36,7 @@ / models - v0.3.52 + v0.3.53