Issue authored by Claude (Anthropic agent) on behalf of @ohdearquant.
Two small per-prefill costs in CPU batched prefill.
- Final RMSNorm on every row.
prefill_prompt in crates/inference/src/forward/batch_prefill.rs normalizes all seq_len rows of scratch.hidden, then reads only the last row (let last_hidden = &scratch.hidden[(seq_len - 1) * hidden..]). Nothing else reads the buffer: PrefillScratch is private and per call, and prefill_prompt returns only the logits. qwen35_rms_norm is per-row, so normalizing just the last row gives identical output.
- LoRA no-op calls. The 13 direct
lora.apply( sites in batch_prefill.rs bypass apply_lora_rows, which checks is_active. So with no adapter loaded, every projection still makes a dynamic call through Box<dyn LoraHook>. That is about 186·T calls per prefill.
Both are small next to attention and GEMM work; they are filed as cheap cleanups. Reach: CPU serve and generate prefill, dense configs. Measurement lane: qwen35_generate TTFT.
Two small per-prefill costs in CPU batched prefill.
prefill_promptincrates/inference/src/forward/batch_prefill.rsnormalizes allseq_lenrows ofscratch.hidden, then reads only the last row (let last_hidden = &scratch.hidden[(seq_len - 1) * hidden..]). Nothing else reads the buffer:PrefillScratchis private and per call, andprefill_promptreturns only the logits.qwen35_rms_normis per-row, so normalizing just the last row gives identical output.lora.apply(sites inbatch_prefill.rsbypassapply_lora_rows, which checksis_active. So with no adapter loaded, every projection still makes a dynamic call throughBox<dyn LoraHook>. That is about 186·T calls per prefill.Both are small next to attention and GEMM work; they are filed as cheap cleanups. Reach: CPU serve and generate prefill, dense configs. Measurement lane:
qwen35_generateTTFT.