Skip to content

inference: CPU prefill normalizes every row and calls no-op LoRA hooks #1767

Description

@ohdearquant

Issue authored by Claude (Anthropic agent) on behalf of @ohdearquant.

Two small per-prefill costs in CPU batched prefill.

  1. Final RMSNorm on every row. prefill_prompt in crates/inference/src/forward/batch_prefill.rs normalizes all seq_len rows of scratch.hidden, then reads only the last row (let last_hidden = &scratch.hidden[(seq_len - 1) * hidden..]). Nothing else reads the buffer: PrefillScratch is private and per call, and prefill_prompt returns only the logits. qwen35_rms_norm is per-row, so normalizing just the last row gives identical output.
  2. LoRA no-op calls. The 13 direct lora.apply( sites in batch_prefill.rs bypass apply_lora_rows, which checks is_active. So with no adapter loaded, every projection still makes a dynamic call through Box<dyn LoraHook>. That is about 186·T calls per prefill.

Both are small next to attention and GEMM work; they are filed as cheap cleanups. Reach: CPU serve and generate prefill, dense configs. Measurement lane: qwen35_generate TTFT.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or requestlattice-inferenceAffects the lattice-inference crate (transformer inference)

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions