Repository navigation
G4 passes at 6.184x once the gate and up projections reach the device - #238
Merged
Merged
Conversation
Same model, same prompts, same --arm both protocol as the 2026-10-07-g4-dispatch
baseline, on the same host class. The PTX module is byte-identical across both
runs -- kernel sha256 4c534e62... either side -- so nothing in the device code
moved and the comparison isolates the routing.
baseline dual path
accelerated decode 9.02 tok/s 31.26 tok/s +246.6%
control decode 5.05 tok/s 5.06 tok/s +0.2%
G4 1.788x FAIL 6.184x PASS gate 3.00x
qualified false true
launches / decode step 241.0 321.0 +33.2%
transfers / decode step 402.0 522.0 +29.9%
Q4_K projections per layer 4.05 6.07 predicted 6.00
The control arm moving 0.2% is what makes the rest readable: the denominator is
stable, so the accelerated arm's 3.47x is not a quieter host.
The dispatch cost rose and it did not matter. 33% more launches and 30% more
transfers bought a 3.47x faster arm, because the two projections that moved are
53.3% of a layer's projection arithmetic. Reading 1.788x as a dispatch ceiling
was wrong -- dispatch was not the binding constraint there, an unimplemented
code path was. Q4_K/DECODE_PROJECTION per layer landing on 6.07 against the 6.00
predicted from the GGUF header confirms the dual path is taken, independent of
any timing.
Correctness is not the claim. TensorOps.ggufDualMatmul computed those
projections on the CPU and produced right answers; it was slower, not wrong.
This change has no correctness credit and would have earned nothing had the
speed not moved. It is kept because 31.26 beats 9.02, and for no other reason.
Also fixes the half of the counter work this run exposed as useless.
CudaRoutingCounters.declined was added and incremented, but
CudaKernelGateCli.Routing never serialised it, so the run printed
totalDeclinedProjections = None: the counter built to make a silent fallback
visible was itself invisible in the artifact. Routing now carries
declinedProjections and totalDeclinedProjections, verified off-device against a
real report rather than assumed.
Still open, and the number that reprioritises it is here: 522 transfers per
decode step moving 15.4 MB remains the standing overhead, but now sits on an arm
passing its gate at 2.06x the threshold, so device-resident activations should
be weighed against G1 rather than assumed next. G1 itself is untouched -- this
run reports tokenParity: null and says nothing about the attention argmax, so
the path passes G4 and still cannot ship.
:backend-cuda:check and :models-bench:check green: 50 and 269 tests, zero
failures.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Same model, same prompts, same
--arm bothprotocol as the2026-10-07-g4-dispatchbaseline. The PTX module is byte-identical across both runs (kernel sha2564c534e62…either side), so nothing in the device code moved and the comparison isolates the routing.qualifiedThe control arm moving 0.2% is what makes the rest readable: the denominator is stable, so the accelerated arm's 3.47x is not a quieter host.
The dispatch cost rose and it did not matter
33% more launches and 30% more transfers bought a 3.47x faster arm, because the two projections that moved are 53.3% of a layer's projection arithmetic. Reading 1.788x as a dispatch ceiling was wrong — dispatch was not the binding constraint there, an unimplemented code path was.
Q4_K/DECODE_PROJECTIONper layer landing on 6.07 against the 6.00 predicted from the GGUF header confirms the dual path is taken, independent of any timing.Correctness is not the claim
TensorOps.ggufDualMatmulcomputed those projections on the CPU and produced right answers — slower, not wrong. This change has no correctness credit and would have earned nothing had the speed not moved. It is kept because 31.26 beats 9.02, and for no other reason.Also fixes the half of #237 this run exposed as useless
CudaRoutingCounters.declinedwas added and incremented, butCudaKernelGateCli.Routingnever serialised it — so the run printedtotalDeclinedProjections = None. The counter built to make a silent fallback visible was itself invisible in the artifact.Routingnow carriesdeclinedProjectionsandtotalDeclinedProjections, verified off-device against a real report.Still open
522 transfers per decode step moving 15.4 MB remains the standing overhead — but it now sits on an arm passing its gate at 2.06x the threshold, so device-resident activations should be weighed against G1 rather than assumed next.
G1 is untouched. This run reports
tokenParity: nulland says nothing about the attention argmax. The path passes G4 and still cannot ship.:backend-cuda:checkand:models-bench:checkgreen: 50 and 269 tests, zero failures.