refactor(inference): build the Metal serving worker through a runtime factory - #1789
Merged
Merged
Conversation
… factory The Metal serving worker took a loader closure that returned `MetalQwen35State` and ran Qwen-specific generation, LoRA residency and vision handling inline in the worker thread. Move that model work behind a crate-private, non-Send `ServingRuntime`, built on the worker thread by a one-shot `RuntimeFactory`. The worker loop keeps its queue, admission, cancellation and terminal-event logic and calls the runtime for generation and adapter control. `ServingFactory::qwen_metal` wraps the existing Qwen loader. Both serving binaries and the `bench_serve_prepare` example build their Metal worker through it. The factory also returns a `PreparationHandle` bound to the worker's tokenizer and context window; the handlers use it for the same request preparation they ran before. CPU serving is unchanged. No behaviour change is intended: refusal codes, window policies, streaming, cancellation, prefix reuse, LoRA application and vision are preserved.
… boundary baseline
E2E Parity ReportPASS: all 4 prompts match within their respective match windows
|
print(fib
|
E2E Parity ReportPASS: 3/4 gating prompts match; 1 known divergence (#535) excluded from the verdict
|
|
…s a test helper in the serve-preparation example
Q4 perplexity regression gate (#616)
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
ADR-090 R07, first PR: the Metal serving worker builds its runtime through a model-owned factory instead of each binary handing it a concrete Qwen state.
Files changed: 13 · commits: 4
Change
serving_factory::ServingFactory(new,src/serving_factory.rs,#[doc(hidden)], compiled only withserveandmetal-gpuon macOS). It sits at the crate root, outsidesrc/model/,src/serve/and the binaries.ServingFactory::qwen_metal(loader, vision)wraps the existing loader closure.buildruns on the worker thread and returns a boxed runtime, the worker metadata and aPreparationHandle. This PR is Qwen-only: it reads nomodel_type, adds no Gemma arm, and keeps the legacy Qwen admission and startup-error behaviour.model::serving_runtime(new, crate-private): an object-safeServingRuntimetrait and its Qwen Metal implementation,QwenMetalRuntime, which carries the per-job generation (including LoRA application and vision), adapter load/unload control and the vision-support flag that used to live inline in the worker. The trait is deliberately notSend: the runtime is built and used on the worker thread only, and a static assertion inserving_runtime/mod.rsfails the build ifSendis ever added.serve::prepare::PreparationHandle(#[doc(hidden)]): request preparation bound to the same tokenizer and context window as the runtime the factory built.prepare_latticechecks the context window with the handle's own tokenizer, on the untruncated prompt length, before stop parsing.MetalWorker::spawn_with_visiontakes aServingFactory. Thelattice serveMetal path,lattice_serveand thebench_serve_prepareexample now go through it, and the two binaries and the example prepare requests through the returned handle. The CPU serve path is unchanged.MetalQwen35Stateconstruction site is updated (src/bin/lattice/serve.rs:393:85).RequestedChatOptionsdisposition (ADR-090, the R07 item): it stays#[doc(hidden)]and internal (serve/contract.rs,serve/prompt_adapter.rs). No binary or example constructs it, and the binaries reach Qwen preparation only throughPreparationHandle. It is not yet folded into the factory. Gemma preparation (prepare_gemma_chat_request) keeps its hiddenpubstatus for the measurement example until Gemma is routed through the factory (R08).Pipeline boundary baseline
Unchanged. The boundary scan treats anything defined under
src/model/orforward/metal_*as model-owned. The factory lives outside both, so naming it from the binaries and the worker adds no crossing. The binaries still constructMetalQwen35Stateinside their loader closures, because this PR does not move loading. Those concrete crossings stay until loading moves behind the factory.Tests and must-fail controls
Each prediction was written before its run. Every mutation followed the same cycle:
The groups follow the change's acceptance list: confinement, admission and cancellation, preparation order, runtime features, and production entry points. Every row not marked as an exception had green unmodified and restored runs and a mutated run that failed as predicted. The admission and two handle-preparation rows ran on the commit before the factory move. That move changed only one import line in
metal_worker.rsand leftprepare.rsuntouched. Every other row ran at the final head.trait ServingRuntime: SendE0283from the not-SendassertionE0283. An earlier run predicted "cannot be sent between threads" and gotE0283; the prediction was corrected.submit_rejects_once_admission_cap_reachedfailsadmission_slot_is_released_when_a_queued_job_is_cancelledrunning_job_cancelled_during_prefill_like_phase_never_calls_on_tokenfailsstopbefore the context checkcm_serve_context_window_checked_before_stop_parsingfailsinvalid_stopin place ofcontext_length_exceededserve_chat_completions_reaches_vision_forward_pathfailslattice servehands the worker the old(loader, vision)pairServingFactoryTwo further mutations named in the acceptance list were not run:
Neither has a text-level mutation that compiles.
Local gate (macOS,
serve,metal-gpu,f16unless noted)cargo fmt --all -- --checkrc 0cargo clippy -p lattice-inference --all-targets --features serve,metal-gpu,f16 -- -D warningsrc 0. Also clippy rc 0 with default features,serve,serve,f16andserve,metal-gpu,f16.--lib serving_factory::: 1 passed--lib serve::: 258 passed--test metal_measurement_lock_contract: 65 passed--bin lattice -- serve: 79 passed--bin lattice_serve: 69 passed--test pipeline_boundary_contract: 14 passed, with default features and withmetal-gpu,f16--test mmap_trust_boundary_contract: 3 passed--test serve_prepare_records: 1 passed, 1 ignored--test vision_serve_e2e_test: 2 passedmacOS feature-matrix steps, re-run on an Apple silicon Mac mini at the final head:
cargo fmt --all -- --check, rc 0.-D warnings, rc 0 in all four configurations the job uses:--release --all-targets --features f16,metal-gpu,bench-internals;--all-targets --features f16,metal-gpu;--all-targets(default features);--all-targets --features mixture.cargo check --workspace --features metal-gpu, rc 0.--example bench_serve_preparetests: 6 passed with default features, and 6 passed withf16,metal-gpu,bench-internals.--bin lattice_serve131 passed,--bin lattice196 passed (both withf16,metal-gpu,test-utils),--bin chat_metal14 passed.pipeline_boundary_contractwithmetal-gpu,f16: 14 passed.The release clippy run with
bench-internalsfailed on the previous head. That flag set compiles the example's real Metal route, and the oldlattice_serve_preparehelper was no longer called from it. The helper now stays compiled for the example's tests only, and dead code is allowed outside tests. The example's route docs now name the handle methods.--no-default-featuresfails to build indownload.rs. That failure is identical on the base commit and is not touched here.bench-compare disposition
Reachable target:
bench_serve_prepare(example,f16,metal-gpu,serve,bench-internals) drives the Metal worker on both serving routes, which this change reaches directly. The target is a plainfn main()and has no Criterion baseline, somake bench-comparecannot pair it. It was paired by hand instead.5e02c363vs head77692001. The two later commits change no timed code. One records a baseline line and the other moves the factory module, and together they touch only import paths and the boundary baseline file. So the PAIR was not re-run.base₁ head₁ head₂ base₂) with the arms adjacent, 5 timed runs per case after a warm-up.First-token latency (ms), median per arm (base / head):
Admission time is 0.001 ms in every arm.
Reading: no measurable change in time to first token on either route. The
lattice-metalpooled numbers read slightly faster at head only because thebase₂arm is about 0.6 ms slow. The idle probe before that arm passed at 75.3%, with a Spotlight indexer (corespotlightd) at 260% CPU.base₁and both head arms agree within about 0.1 ms. This looks like contamination of one arm, not an effect of the change.Other timings in the run:
load_mshas one sample per arm. The first arm of each route includes a cold file cache, so no conclusion is drawn from it.prepare_msis not the same code in the two arms. Base timed the example's local copy of the preparation logic, while head times the productionPreparationHandlepath. Both stay at hundredths of a millisecond.Residual risk, unmeasured: the vision route and LoRA-carrying requests were not benched.