Skip to content

refactor(inference): build the Metal serving worker through a runtime factory - #1789

Merged
ohdearquant merged 4 commits into
mainfrom
refactor/r07-serving-factory
Sep 25, 2026
Merged

ohdearquant merged 4 commits into
mainfrom
refactor/r07-serving-factory

Conversation

@ohdearquant

@ohdearquant ohdearquant commented Sep 25, 2026 •

Copy link
Copy Markdown
Owner

PR description authored by Claude (Anthropic agent) on behalf of @ohdearquant.

ADR-090 R07, first PR: the Metal serving worker builds its runtime through a model-owned factory instead of each binary handing it a concrete Qwen state.

Files changed: 13 · commits: 4

Change

  • serving_factory::ServingFactory (new, src/serving_factory.rs, #[doc(hidden)], compiled only with serve and metal-gpu on macOS). It sits at the crate root, outside src/model/, src/serve/ and the binaries. ServingFactory::qwen_metal(loader, vision) wraps the existing loader closure. build runs on the worker thread and returns a boxed runtime, the worker metadata and a PreparationHandle. This PR is Qwen-only: it reads no model_type, adds no Gemma arm, and keeps the legacy Qwen admission and startup-error behaviour.
  • model::serving_runtime (new, crate-private): an object-safe ServingRuntime trait and its Qwen Metal implementation, QwenMetalRuntime, which carries the per-job generation (including LoRA application and vision), adapter load/unload control and the vision-support flag that used to live inline in the worker. The trait is deliberately not Send: the runtime is built and used on the worker thread only, and a static assertion in serving_runtime/mod.rs fails the build if Send is ever added.
  • serve::prepare::PreparationHandle (#[doc(hidden)]): request preparation bound to the same tokenizer and context window as the runtime the factory built. prepare_lattice checks the context window with the handle's own tokenizer, on the untruncated prompt length, before stop parsing.
  • MetalWorker::spawn_with_vision takes a ServingFactory. The lattice serve Metal path, lattice_serve and the bench_serve_prepare example now go through it, and the two binaries and the example prepare requests through the returned handle. The CPU serve path is unchanged.
  • The Metal lock contract's recorded position for the moved MetalQwen35State construction site is updated (src/bin/lattice/serve.rs:393:85).

RequestedChatOptions disposition (ADR-090, the R07 item): it stays #[doc(hidden)] and internal (serve/contract.rs, serve/prompt_adapter.rs). No binary or example constructs it, and the binaries reach Qwen preparation only through PreparationHandle. It is not yet folded into the factory. Gemma preparation (prepare_gemma_chat_request) keeps its hidden pub status for the measurement example until Gemma is routed through the factory (R08).

Pipeline boundary baseline

Unchanged. The boundary scan treats anything defined under src/model/ or forward/metal_* as model-owned. The factory lives outside both, so naming it from the binaries and the worker adds no crossing. The binaries still construct MetalQwen35State inside their loader closures, because this PR does not move loading. Those concrete crossings stay until loading moves behind the factory.

Tests and must-fail controls

Each prediction was written before its run. Every mutation followed the same cycle:

  1. Run the named tests on unmodified source (green).
  2. Apply the mutation, touch the file and re-run.
  3. Restore the source and re-run (green again).

The groups follow the change's acceptance list: confinement, admission and cancellation, preparation order, runtime features, and production entry points. Every row not marked as an exception had green unmodified and restored runs and a mutated run that failed as predicted. The admission and two handle-preparation rows ran on the commit before the factory move. That move changed only one import line in metal_worker.rs and left prepare.rs untouched. Every other row ran at the final head.

Group Mutation Predicted Mutated result
Confinement trait ServingRuntime: Send lib build fails, E0283 from the not-Send assertion fails with E0283. An earlier run predicted "cannot be sent between threads" and got E0283; the prediction was corrected.
Admission add one permit back when a job is dequeued submit_rejects_once_admission_cap_reached fails failed
Cancellation skip the dequeue-time cancel check both dequeue-cancel tests fail failed, plus admission_slot_is_released_when_a_queued_job_is_cancelled
Cancellation drop the cancel flag from the in-flight cancel predicate running_job_cancelled_during_prefill_like_phase_never_calls_on_token fails failed
Cancellation drop the closed-receiver check from the in-flight predicate predicted no test fails (coverage gap) none of the 42 worker tests failed: gap confirmed, not introduced here; the code is unchanged from before this PR (#1791)
Preparation order parse stop before the context check cm_serve_context_window_checked_before_stop_parsing fails failed
Preparation handle's tokenizer count divided by 16 new prepare test fails at the over-context arm failed with invalid_stop in place of context_length_exceeded
Preparation handle's window replaced by 4096 same same
Runtime: shutdown unbounded join in place of the deadline wait the two deadline tests hang or fail (run under a hard timeout) both failed
Runtime: vision runtime reports vision unsupported real-checkpoint serve_chat_completions_reaches_vision_forward_path fails failed: the server answered the image request with an HTTP error status. Run on an Apple silicon Mac mini with the real Qwen3.5-0.8B checkpoint; the unmodified arm passed 2 of 2. No separate restored run: the mutation was applied to a fresh checkout of the head, so the unmodified arm ran on the restored bytes.
Runtime: prefix reuse, LoRA application, control acknowledgement order none run none no automated control. These paths need a real model and no existing test exercises them through the worker. LoRA coverage is tracked in #1790.
Entry points lattice serve hands the worker the old (loader, vision) pair bin build fails, expecting ServingFactory failed as predicted

Two further mutations named in the acceptance list were not run:

  • constructing Metal state before the worker thread spawns;
  • replacing admission refusal with waiting for capacity.

Neither has a text-level mutation that compiles.

Local gate (macOS, serve,metal-gpu,f16 unless noted)

  • cargo fmt --all -- --check rc 0
  • cargo clippy -p lattice-inference --all-targets --features serve,metal-gpu,f16 -- -D warnings rc 0. Also clippy rc 0 with default features, serve, serve,f16 and serve,metal-gpu,f16.
  • --lib serving_factory::: 1 passed
  • --lib serve::: 258 passed
  • --test metal_measurement_lock_contract: 65 passed
  • --bin lattice -- serve: 79 passed
  • --bin lattice_serve: 69 passed
  • --test pipeline_boundary_contract: 14 passed, with default features and with metal-gpu,f16
  • --test mmap_trust_boundary_contract: 3 passed
  • --test serve_prepare_records: 1 passed, 1 ignored
  • --test vision_serve_e2e_test: 2 passed

macOS feature-matrix steps, re-run on an Apple silicon Mac mini at the final head:

  • cargo fmt --all -- --check, rc 0.
  • Clippy with -D warnings, rc 0 in all four configurations the job uses:
    • --release --all-targets --features f16,metal-gpu,bench-internals;
    • --all-targets --features f16,metal-gpu;
    • --all-targets (default features);
    • --all-targets --features mixture.
  • cargo check --workspace --features metal-gpu, rc 0.
  • --example bench_serve_prepare tests: 6 passed with default features, and 6 passed with f16,metal-gpu,bench-internals.
  • Binary tests: --bin lattice_serve 131 passed, --bin lattice 196 passed (both with f16,metal-gpu,test-utils), --bin chat_metal 14 passed.
  • pipeline_boundary_contract with metal-gpu,f16: 14 passed.

The release clippy run with bench-internals failed on the previous head. That flag set compiles the example's real Metal route, and the old lattice_serve_prepare helper was no longer called from it. The helper now stays compiled for the example's tests only, and dead code is allowed outside tests. The example's route docs now name the handle methods.

--no-default-features fails to build in download.rs. That failure is identical on the base commit and is not touched here.

bench-compare disposition

Reachable target: bench_serve_prepare (example, f16,metal-gpu,serve,bench-internals) drives the Metal worker on both serving routes, which this change reaches directly. The target is a plain fn main() and has no Criterion baseline, so make bench-compare cannot pair it. It was paired by hand instead.

  • Setup: an Apple silicon Mac mini, Qwen3.5-0.8B Q4 checkpoint, base 5e02c363 vs head 77692001. The two later commits change no timed code. One records a baseline line and the other moves the factory module, and together they touch only import paths and the boundary baseline file. So the PAIR was not re-run.
  • Order: each route ran ABBA (base₁ head₁ head₂ base₂) with the arms adjacent, 5 timed runs per case after a warm-up.
  • Machine: the whole run held the machine bench window, and a CPU-idle probe (floor 70%) passed before every arm.
  • Not covered: there was no in-phase sampling, and no cooldown, AC-power, thermal or HID-idle checkpoint.

First-token latency (ms), median per arm (base / head):

Route Case base₁ head₁ head₂ base₂ Pooled base → head
lattice-serve control 49.93 49.80 49.78 49.81 49.81 → 49.79 (−0.05%)
lattice-serve multi_turn 93.99 94.17 93.84 93.72 93.73 → 94.11 (+0.40%)
lattice-serve reasoning 49.60 49.64 49.65 49.53 49.53 → 49.64 (+0.22%)
lattice-metal control 49.17 49.09 48.87 49.76 49.24 → 48.98 (−0.54%)
lattice-metal multi_turn 93.11 93.23 93.11 93.92 93.52 → 93.16 (−0.39%)
lattice-metal reasoning 48.99 49.03 49.01 49.64 49.50 → 49.02 (−0.97%)

Admission time is 0.001 ms in every arm.

Reading: no measurable change in time to first token on either route. The lattice-metal pooled numbers read slightly faster at head only because the base₂ arm is about 0.6 ms slow. The idle probe before that arm passed at 75.3%, with a Spotlight indexer (corespotlightd) at 260% CPU. base₁ and both head arms agree within about 0.1 ms. This looks like contamination of one arm, not an effect of the change.

Other timings in the run:

  • load_ms has one sample per arm. The first arm of each route includes a cold file cache, so no conclusion is drawn from it.
  • prepare_ms is not the same code in the two arms. Base timed the example's local copy of the preparation logic, while head times the production PreparationHandle path. Both stay at hundredths of a millisecond.

Residual risk, unmeasured: the vision route and LoRA-carrying requests were not benched.

… factory

The Metal serving worker took a loader closure that returned
`MetalQwen35State` and ran Qwen-specific generation, LoRA residency and
vision handling inline in the worker thread. Move that model work behind a
crate-private, non-Send `ServingRuntime`, built on the worker thread by a
one-shot `RuntimeFactory`. The worker loop keeps its queue, admission,
cancellation and terminal-event logic and calls the runtime for generation
and adapter control.

`ServingFactory::qwen_metal` wraps the existing Qwen loader. Both serving
binaries and the `bench_serve_prepare` example build their Metal worker
through it. The factory also returns a `PreparationHandle` bound to the
worker's tokenizer and context window; the handlers use it for the same
request preparation they ran before. CPU serving is unchanged.

No behaviour change is intended: refusal codes, window policies, streaming,
cancellation, prefix reuse, LoRA application and vision are preserved.
@github-actions

github-actions Bot commented Sep 25, 2026 •

Copy link
Copy Markdown

E2E Parity Report

PASS: all 4 prompts match within their respective match windows

Prompt Window Agreement First Diff HF tok/s Lattice tok/s Verdict
The capital of France is 3 15/15 none frozen 1.3 PASS
In the year 2024, artificial intelligence 3 15/15 none frozen 1.1 PASS
`def fibonacci(n):
if n <= 1:
    return n
return` | 3 | 15/15 | none | frozen | 1.1 | PASS |

| def merge_sort(arr): """ Merge sort implementation. | 2 | 15/15 | none | frozen | 0.5 | PASS |

The capital of France is

  • HF: frozen token ids: [11751, 13, 198, 760, 6511, 314, 9338, 369, 11751, 13, 198, 76
  • Lattice: Paris.
    The capital of France is Paris.
    The capital of France

In the year 2024, artificial intelligence

  • HF: frozen token ids: [318, 15015, 8, 682, 3512, 264, 4927, 919, 314, 279, 3521, 832
  • Lattice: (AI) has become a significant part of the global economy. It is

def fibonacci(n): if n <= 1: return n return

  • HF: frozen token ids: [73111, 1393, 12, 16, 8, 478, 73111, 1393, 12, 17, 8, 271, 130
  • Lattice: fibonacci(n-1) + fibonacci(n-2)

print(fib

def merge_sort(arr): """ Merge sort implementation.

  • HF: frozen token ids: [10562, 17885, 10620, 1590, 198, 262, 4071, 198, 262, 38754, 3
  • Lattice: merge_sort(arr):
    """
    Merge sort implementation.

@github-actions

github-actions Bot commented Sep 25, 2026 •

Copy link
Copy Markdown

E2E Parity Report

PASS: 3/4 gating prompts match; 1 known divergence (#535) excluded from the verdict

Prompt Window Agreement First Diff HF tok/s Lattice tok/s Verdict
The capital of France is 3 4/4 none frozen 6.5 PASS
In the year 2024, artificial intelligence 3 4/4 none frozen 6.9 PASS
`def fibonacci(n):
if n <= 1:
    return n
return` | 3 | 4/4 | none | frozen | 6.4 | PASS |

| def merge_sort(arr): """ Merge sort implementation. | 2 | 0/4 | pos 0 | frozen | 0.8 | KNOWN-DIVERGENT (#535) |

The capital of France is

  • HF: frozen token ids: [11751, 13, 198, 760]
  • Lattice: Paris.
    The

In the year 2024, artificial intelligence

  • HF: frozen token ids: [318, 15015, 8, 682]
  • Lattice: (AI) has

def fibonacci(n): if n <= 1: return n return

  • HF: frozen token ids: [73111, 1393, 12, 16]
  • Lattice: fibonacci(n-1

def merge_sort(arr): """ Merge sort implementation.

  • HF: frozen token ids: [10562, 17885, 10620, 1590]
  • Lattice: main():

…s a test helper in the serve-preparation example
@github-actions

github-actions Bot commented Sep 25, 2026 •

Copy link
Copy Markdown

Q4 perplexity regression gate (#616)

  • tier: unrotated Metal Q4 (shipping)
  • corpus: docs/bench_results/wiki.test.raw, first 2048 tokens
  • window/stride: 512/256
  • tokens scored: 2047
  • measured PPL: 16.589109
  • golden PPL: 16.589111
  • tolerance: 0.050000
  • |measured - golden|: 0.000002
  • verdict: PASS

@ohdearquant
ohdearquant merged commit a243592 into main Sep 25, 2026
51 checks passed
@ohdearquant
ohdearquant deleted the refactor/r07-serving-factory branch September 25, 2026 18:21
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant