Skip to content

fix(inference): keep long prompts whole in lattice chat - #1793

Merged
ohdearquant merged 1 commit into
mainfrom
fix/chat-prompt-truncation
Sep 25, 2026
Merged

ohdearquant merged 1 commit into
mainfrom
fix/chat-prompt-truncation

Conversation

@ohdearquant

Copy link
Copy Markdown
Owner

PR description authored by Claude (Anthropic agent) on behalf of @ohdearquant.

Fixes #1788.

What was wrong

lattice chat tokenized each line with the model's default 4096-token tokenizer cap.

  • CPU backend: Qwen35Model::max_context() is min(max_position_embeddings, 8192), but generation re-tokenizes with the capped tokenizer and keeps input_ids[..real_length]. A prompt of 4097..max_context tokens was generated from its first 4096 tokens, and the printed [N prompt tokens, …] showed 4096.
  • Metal backend: the chat KV cache is 4096 (MAX_CACHE_LEN). An over-long prompt was rejected by the context-budget check, but with a count taken from the already-shortened prompt. With a zero decode budget it was returned as a shortened, empty generation.

Change

All runtime changes are in src/bin/lattice/chat.rs:

  • Backend::cpu raises the CPU model's tokenizer cap to max_context() at load, using Qwen35Model::ensure_tokenizer_max_seq_len (added in fix(serve): keep long chat prompts whole on the CPU serving path #1787). This matches lattice serve's CPU backend.
  • generate_checked counts the prompt's full length (pre_truncation_len from the tokenizer the backend generates with). Before calling generation, it refuses a prompt longer than the backend's context window with InvalidInput("prompt (N tokens) exceeds model context window (LIMIT)"). Both the CPU and Metal backends go through it.
  • The per-line dispatch moves into Backend::generate_chat_line. The REPL still passes the trimmed line, prints the error, and reads the next line.

tests/data/pipeline_boundary_baseline.txt gains one row for the test module's use of test_support::tiny_zero_model_with_context, the same fixture bin/lattice/main.rs tests already use.

Not changed: src/bin/chat_metal.rs is a separate executable with the same 4096 tokenizer/cache pairing. Its prefix-cache path reports shortened counts the same way the Metal backend here did. It is left for its own change.

Tests

New tests in chat.rs, under cfg(all(test, feature = "test-utils")):

Test Observes
cpu_chat_generation_keeps_long_prompts GenerateOutput.prompt_tokens from real CPU generation for 4097 and 8192 tokens, with context 8192 and initial cap 4096
cpu_chat_refuses_full_count_and_accepts_next_line limit+1 and limit+137 refused with the exact message at limits 32 and 8192; the next short line is accepted
chat_guard_refuses_before_generation_with_metal_tokenizer_cap at a 4096 cap and limit, 4097/4233 are refused before the generation callback runs; 1 and 4096 pass through unchanged
metal_chat_uses_checked_generation (metal-gpu) source-level check that the Metal wrapper routes through the guard with its own tokenizer and max_context()
repl_uses_checked_generation_and_cpu_initialization source-level check that the stdin loop uses Backend::cpu and generate_chat_line(trimmed, …)

The two source-level tests guard wiring that the other tests can't reach. The REPL reads stdin, and the Metal fixtures that fit the flash-decode shape are private to the library's own tests. No real Metal load or generation runs in this PR's tests.

Each test was checked for mutation sensitivity on an Apple silicon Mac mini. Each fix line was removed in turn (then touch), the named test run, and the line restored:

Mutation removed Test fixed mutant restored
CPU tokenizer cap cpu_chat_generation_keeps_long_prompts ok FAILED ok
guard refusal chat_guard_refuses_before_generation_with_metal_tokenizer_cap ok FAILED ok
full count (pre_truncation_len → real_length) chat_guard_refuses_before_generation_with_metal_tokenizer_cap ok FAILED ok
CPU guard wiring cpu_chat_refuses_full_count_and_accepts_next_line ok FAILED ok
Metal guard wiring metal_chat_uses_checked_generation ok FAILED ok
CPU load wiring repl_uses_checked_generation_and_cpu_initialization ok FAILED ok

Gate (Apple silicon Mac mini, at this head's tree on base a243592)

  • cargo fmt --all -- --check rc 0
  • clippy -D warnings, -p lattice-inference --all-targets, rc 0 with each feature set: test-utils; f16,metal-gpu,test-utils; --release f16,metal-gpu,bench-internals; f16,metal-gpu; default; mixture
  • cargo test -p lattice-inference --bin lattice --features test-utils: 191 passed
  • cargo test -p lattice-inference --bin lattice --features f16,metal-gpu,test-utils: 203 passed
  • pipeline_boundary_contract: 14 passed at default features and at metal-gpu,f16

bench-compare disposition

No bench run: this is a structural waiver. The runtime change is confined to the lattice binary's interactive chat subcommand (src/bin/lattice/chat.rs), which is reached only through Command::Chat in bin/lattice/main.rs. I searched every declared target in every crate manifest: 25 bench targets and 14 bin targets in crates/inference, 5 benches and 2 bins in crates/embed, 1 bench each in crates/fann and crates/transport, and 4 bins in crates/tune. None of them calls the chat module, generate_chat_line or generate_checked. The other changed file is test baseline data. Manifests, lockfile, features, profiles and bench harnesses are unchanged. Residual risk: the added per-line tokenization for the length check is unmeasured. It runs once per interactive line, outside any timed region.

Files changed: 2 · commits: 1

`lattice chat` tokenized each line with the model's default 4096-token
tokenizer cap, so on the CPU backend a prompt between 4097 tokens and the
model's context window was generated from its first 4096 tokens, and the
printed prompt count showed the shortened length. On the Metal backend an
over-long prompt was rejected with a count taken from the already-shortened
prompt, and with a zero decode budget it was shortened silently.

The CPU backend now raises the tokenizer cap to the model's context window
when it loads, as `lattice serve` does. Before generating, both backends
count the prompt's full length and refuse a prompt longer than the context
window with an error naming its real token count and the limit; the REPL
reports the error and reads the next line.
@github-actions

Copy link
Copy Markdown

E2E Parity Report

PASS: 3/4 gating prompts match; 1 known divergence (#535) excluded from the verdict

Prompt Window Agreement First Diff HF tok/s Lattice tok/s Verdict
The capital of France is 3 4/4 none frozen 0.8 PASS
In the year 2024, artificial intelligence 3 4/4 none frozen 4.5 PASS
`def fibonacci(n):
if n <= 1:
    return n
return` | 3 | 4/4 | none | frozen | 2.0 | PASS |

| def merge_sort(arr): """ Merge sort implementation. | 2 | 0/4 | pos 0 | frozen | 0.7 | KNOWN-DIVERGENT (#535) |

The capital of France is

  • HF: frozen token ids: [11751, 13, 198, 760]
  • Lattice: Paris.
    The

In the year 2024, artificial intelligence

  • HF: frozen token ids: [318, 15015, 8, 682]
  • Lattice: (AI) has

def fibonacci(n): if n <= 1: return n return

  • HF: frozen token ids: [73111, 1393, 12, 16]
  • Lattice: fibonacci(n-1

def merge_sort(arr): """ Merge sort implementation.

  • HF: frozen token ids: [10562, 17885, 10620, 1590]
  • Lattice: main():

@github-actions

Copy link
Copy Markdown

E2E Parity Report

PASS: all 4 prompts match within their respective match windows

Prompt Window Agreement First Diff HF tok/s Lattice tok/s Verdict
The capital of France is 3 15/15 none frozen 0.6 PASS
In the year 2024, artificial intelligence 3 15/15 none frozen 0.6 PASS
`def fibonacci(n):
if n <= 1:
    return n
return` | 3 | 15/15 | none | frozen | 0.8 | PASS |

| def merge_sort(arr): """ Merge sort implementation. | 2 | 15/15 | none | frozen | 0.5 | PASS |

The capital of France is

  • HF: frozen token ids: [11751, 13, 198, 760, 6511, 314, 9338, 369, 11751, 13, 198, 76
  • Lattice: Paris.
    The capital of France is Paris.
    The capital of France

In the year 2024, artificial intelligence

  • HF: frozen token ids: [318, 15015, 8, 682, 3512, 264, 4927, 919, 314, 279, 3521, 832
  • Lattice: (AI) has become a significant part of the global economy. It is

def fibonacci(n): if n <= 1: return n return

  • HF: frozen token ids: [73111, 1393, 12, 16, 8, 478, 73111, 1393, 12, 17, 8, 271, 130
  • Lattice: fibonacci(n-1) + fibonacci(n-2)

print(fib

def merge_sort(arr): """ Merge sort implementation.

  • HF: frozen token ids: [10562, 17885, 10620, 1590, 198, 262, 4071, 198, 262, 38754, 3
  • Lattice: merge_sort(arr):
    """
    Merge sort implementation.

@github-actions

Copy link
Copy Markdown

Q4 perplexity regression gate (#616)

  • tier: unrotated Metal Q4 (shipping)
  • corpus: docs/bench_results/wiki.test.raw, first 2048 tokens
  • window/stride: 512/256
  • tokens scored: 2047
  • measured PPL: 16.589109
  • golden PPL: 16.589111
  • tolerance: 0.050000
  • |measured - golden|: 0.000002
  • verdict: PASS

@ohdearquant
ohdearquant merged commit 17cdea1 into main Sep 25, 2026
50 checks passed
@ohdearquant
ohdearquant deleted the fix/chat-prompt-truncation branch September 25, 2026 19:50
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

chat: lattice chat truncates prompts over 4096 tokens without an error

1 participant