Skip to content

G4 passes at 6.184x once the gate and up projections reach the device - #238

Merged
bsbodden merged 1 commit into
mainfrom
evidence/g4-dualpath-passes
Oct 7, 2026
Merged

bsbodden merged 1 commit into
mainfrom
evidence/g4-dualpath-passes

Conversation

@bsbodden

@bsbodden bsbodden commented Oct 7, 2026

Copy link
Copy Markdown
Member

Same model, same prompts, same --arm both protocol as the 2026-10-07-g4-dispatch baseline. The PTX module is byte-identical across both runs (kernel sha256 4c534e62… either side), so nothing in the device code moved and the comparison isolates the routing.

baseline dual path routed delta
accelerated decode 9.02 tok/s 31.26 tok/s +246.6%
control decode 5.05 tok/s 5.06 tok/s +0.2%
G4 1.788x — FAILED 6.184x — PASSED gate 3.00x
qualified false true
launches / decode step 241.0 321.0 +33.2%
transfers / decode step 402.0 522.0 +29.9%
Q4_K projections per layer 4.05 6.07 predicted 6.00

The control arm moving 0.2% is what makes the rest readable: the denominator is stable, so the accelerated arm's 3.47x is not a quieter host.

The dispatch cost rose and it did not matter

33% more launches and 30% more transfers bought a 3.47x faster arm, because the two projections that moved are 53.3% of a layer's projection arithmetic. Reading 1.788x as a dispatch ceiling was wrong — dispatch was not the binding constraint there, an unimplemented code path was.

Q4_K/DECODE_PROJECTION per layer landing on 6.07 against the 6.00 predicted from the GGUF header confirms the dual path is taken, independent of any timing.

Correctness is not the claim

TensorOps.ggufDualMatmul computed those projections on the CPU and produced right answers — slower, not wrong. This change has no correctness credit and would have earned nothing had the speed not moved. It is kept because 31.26 beats 9.02, and for no other reason.

Also fixes the half of #237 this run exposed as useless

CudaRoutingCounters.declined was added and incremented, but CudaKernelGateCli.Routing never serialised it — so the run printed totalDeclinedProjections = None. The counter built to make a silent fallback visible was itself invisible in the artifact. Routing now carries declinedProjections and totalDeclinedProjections, verified off-device against a real report.

Still open

522 transfers per decode step moving 15.4 MB remains the standing overhead — but it now sits on an arm passing its gate at 2.06x the threshold, so device-resident activations should be weighed against G1 rather than assumed next.

G1 is untouched. This run reports tokenParity: null and says nothing about the attention argmax. The path passes G4 and still cannot ship.

:backend-cuda:check and :models-bench:check green: 50 and 269 tests, zero failures.

Same model, same prompts, same --arm both protocol as the 2026-10-07-g4-dispatch
baseline, on the same host class. The PTX module is byte-identical across both
runs -- kernel sha256 4c534e62... either side -- so nothing in the device code
moved and the comparison isolates the routing.

                               baseline      dual path
  accelerated decode           9.02 tok/s    31.26 tok/s   +246.6%
  control decode               5.05 tok/s     5.06 tok/s     +0.2%
  G4                           1.788x FAIL    6.184x PASS   gate 3.00x
  qualified                    false          true
  launches / decode step       241.0          321.0         +33.2%
  transfers / decode step      402.0          522.0         +29.9%
  Q4_K projections per layer   4.05           6.07          predicted 6.00

The control arm moving 0.2% is what makes the rest readable: the denominator is
stable, so the accelerated arm's 3.47x is not a quieter host.

The dispatch cost rose and it did not matter. 33% more launches and 30% more
transfers bought a 3.47x faster arm, because the two projections that moved are
53.3% of a layer's projection arithmetic. Reading 1.788x as a dispatch ceiling
was wrong -- dispatch was not the binding constraint there, an unimplemented
code path was. Q4_K/DECODE_PROJECTION per layer landing on 6.07 against the 6.00
predicted from the GGUF header confirms the dual path is taken, independent of
any timing.

Correctness is not the claim. TensorOps.ggufDualMatmul computed those
projections on the CPU and produced right answers; it was slower, not wrong.
This change has no correctness credit and would have earned nothing had the
speed not moved. It is kept because 31.26 beats 9.02, and for no other reason.

Also fixes the half of the counter work this run exposed as useless.
CudaRoutingCounters.declined was added and incremented, but
CudaKernelGateCli.Routing never serialised it, so the run printed
totalDeclinedProjections = None: the counter built to make a silent fallback
visible was itself invisible in the artifact. Routing now carries
declinedProjections and totalDeclinedProjections, verified off-device against a
real report rather than assumed.

Still open, and the number that reprioritises it is here: 522 transfers per
decode step moving 15.4 MB remains the standing overhead, but now sits on an arm
passing its gate at 2.06x the threshold, so device-resident activations should
be weighed against G1 rather than assumed next. G1 itself is untouched -- this
run reports tokenParity: null and says nothing about the attention argmax, so
the path passes G4 and still cannot ship.

:backend-cuda:check and :models-bench:check green: 50 and 269 tests, zero
failures.
@bsbodden
bsbodden merged commit 15661fb into main Oct 7, 2026
5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant