Skip to content

Latest commit

 

History

History
300 lines (261 loc) · 112 KB

File metadata and controls

300 lines (261 loc) · 112 KB

Capability matrix: lattice CLI vs lattice_serve

This document tracks which capabilities exist in each of lattice's two OpenAI-shaped binaries, so new features don't silently land in only one of them. It is the reference #589 asked for.

A naming note before anything else

Same trap as docs/serve-http-api.md and docs/cross-turn-cache.md describe -- read this first:

  • lattice CLI (crates/inference/src/bin/lattice/ -- main.rs dispatch, chat.rs REPL, doctor.rs preflight, serve.rs HTTP API; split from a single lattice.rs file in #834) -- one binary, three subcommands: lattice chat (interactive REPL), lattice serve (an HTTP server, OpenAI-compatible), lattice doctor (preflight check, no model weights loaded). This is the README-documented, actively developed surface. Its serve subcommand is itself an HTTP server -- when a row below says "supported" for this column via HTTP, it means lattice serve's router, not lattice chat. Below, unqualified lattice.rs:NNN line citations predate the #834 split and are stale on file/line but not on subcommand/behavior; re-derive against the split files rather than trusting the number.
  • lattice_serve (crates/inference/src/bin/lattice_serve.rs) -- a separate, standalone binary purpose-built as the HTTP daemon the macOS Lattice Studio app spawns and talks to (introduced PR #435). Metal-GPU-only, its own fixed route list, its own request struct.

Every "supported" claim below cites a file:line. Every "not yet" cites the tracking issue. As of ADR-080 cluster C2 (#786) -- see crates/inference/src/serve/mod.rs for the shared contract (error envelope, finish_reason mapping, max_tokens == 0 rejection, /+/v1/models response bodies, disconnect-cancellation guard) both binaries now build their routers on top of.

Every row below is backed by an executable fixture or an explicit not-testable-yet justification: see the Fixture manifest section, checked by scripts/check-capability-matrix.sh (wired into make lint-docs). A row that names a fixture ID citing a test function that no longer exists fails that check closed (#654).

CLI subcommands (lattice binary)

Subcommand Flags Evidence
chat --model (required), --max-tokens (default 256), --temperature (default 0.7), --tokenizer-dir lattice.rs:22-39
serve --model (required), --host (default 127.0.0.1), --port (default 8080), --max-tokens (default 256, per-request default), --model-id, --tokenizer-dir lattice.rs:40-63
doctor --model (required), --context (optional target length), --tokenizer-dir lattice.rs:64-80

No --lora, --reasoning-budget, --top-k, or --repetition-penalty flags exist on any subcommand -- chat only ever builds a GenerateConfig::default() and then assigns max_new_tokens and temperature, nothing else (bin/lattice/chat.rs:149-151), so top_k/repetition_penalty/seed/grammar/MTP/ reasoning-budget are all fixed at their Default values (qwen35_config.rs:2429-2447: top_k: 50, top_p: 0.9, repetition_penalty: 1.1, enable_mtp: None, grammar: None, reasoning_budget: None) for the whole interactive session -- there is no way to change them without editing source.

HTTP routes

lattice serve (subcommand of the lattice binary)

Complete router (lattice.rs, mod serve, router()):

Router::new()
    .route("/", get(root))
    .route("/health", get(health))
    .route("/v1/models", get(list_models))
    .route("/v1/chat/completions", post(chat_completions))
    .route("/v1/embeddings", post(embeddings))
    .route("/v1/lora", get(lora_list))
    .route("/v1/lora/load", post(lora_load))
    .route("/v1/lora/unload", post(lora_unload))
    .layer(DefaultBodyLimit::max(REQUEST_BODY_LIMIT_BYTES))

Both GET / and GET /v1/models landed in ADR-080 C2 (#786): lattice serve previously had neither route. GET / and /v1/models' response bodies are now built from the shared lattice_inference::serve::root_body/models_list_body helpers so both binaries return byte-identical shapes for the same inputs. POST /v1/embeddings is always installed on the router, but every request 400s with vision_unsupported unless the server was started with --model pointed at a vision-language checkpoint (see docs/serve-http-api.md). No /v1/completions, no admin/metrics route.

lattice_serve (standalone binary)

Complete router (lattice_serve.rs, router(), extracted from run() so it's directly testable via tower::ServiceExt::oneshot):

Router::new()
    .route("/", get(root))
    .route("/health", get(health))
    .route("/v1/models", get(list_models))
    .route("/v1/chat/completions", post(chat_completions))
    .route("/v1/embeddings", post(embeddings))
    .route("/v1/lora", get(lora_list))
    .route("/v1/lora/load", post(lora_load))
    .route("/v1/lora/unload", post(lora_unload))

GET /v1/models returns a single-entry OpenAI model list for the one loaded model, via the same shared models_list_body helper lattice serve now uses. Deliberately does NOT install DefaultBodyLimit: this binary enforces the same REQUEST_BODY_LIMIT_BYTES cap manually inside chat_completions via to_bytes, a documented intentional divergence in mechanism (same limit value, different enforcement path) -- see the cross-binary parity table in serve/mod.rs (CHAT_COMPLETIONS_PARITY_CASES) for the full list of aligned-vs-documented-divergent outcomes.

POST /v1/embeddings (issue #584) is always registered, but only answers requests when the binary was started with --embedding-model <dir> (or LATTICE_SERVE_EMBEDDING_MODEL): without it, a request that otherwise passes content-type, body-size, JSON, and input validation gets HTTP 503 embedding_model_not_loaded rather than the route not existing at all, so its absence is a discoverable server state rather than a bare 404. Content-type/body-size/parse/field validation run first and take precedence -- a request with a non-JSON Content-Type, an oversized body, malformed JSON, or an invalid/unsupported field still gets its own 415, 413, or 400 even with no embedding model loaded. The route computes embeddings via lattice_inference::model::bert::BertModel directly (CPU-only), not through lattice-embed's EmbeddingService trait -- lattice-embed already depends on lattice-inference for BertModel/QwenModel, so a reverse edge from this crate back onto lattice-embed is a cyclic package dependency Cargo refuses to build (confirmed with cargo tree -p lattice-inference after adding the edge experimentally: "cyclic package dependency ... lattice-embed ... depends on itself"). An optional --embedding-pooling mean|cls flag selects the pooling strategy (mean is BertModel's own default; BGE-family checkpoints need cls). Request/response DTOs, input normalization (single string or array, untagged), and response building are in the shared, GPU-independent lattice_inference::serve::embeddings module so that logic is unit-testable without a loaded model or the metal-gpu feature; the HTTP handler, CLI flags, and AppState wiring live in lattice_serve.rs itself, gated the same as the rest of this binary's imp module (macOS + metal-gpu feature).

Capability matrix

Capability lattice CLI lattice_serve HTTP
Chat completion (single/multi-turn) supported -- lattice chat (REPL, lattice.rs:2206-2352) and lattice serve POST /v1/chat/completions (chat_completions/chat_completions_with_request, lattice.rs:3616-) supported -- POST /v1/chat/completions (lattice_serve.rs:425-)
Streaming (SSE) supported -- lattice serve only, stream: true (lattice.rs, see docs/serve-http-api.md "streaming"); lattice chat's REPL has no streaming concept (prints the full response once) supported -- stream: true, its own SSE phase machine (lattice_serve.rs:415-)
Disconnect cancellation (mid-stream) supported, both backends -- ADR-080 C2 (#744, #786) wired the shared lattice_inference::serve::CancelOnDrop/cancel_pair guard into the CPU streaming path: cancel_guard is captured by body_stream's flat_map closure so it drops (flipping cancel_rx) the instant axum drops the SSE response on client disconnect, and generate_streaming_with_cancel polls it independently of on_token at four checkpoints (pre-prefill, post-prefill, and the top of every decode-loop iteration, both the fast and stop-string paths). The Metal/Q4 path already had this via the same shared guard threaded through cancel_rx. Previously the CPU path had none (its delta closure ignored the send result, so generation ran to max_tokens after a disconnect) -- that gap is what #744 closed supported -- CancelOnDrop/watch::channel (PR #552/#606), now built on the same shared lattice_inference::serve::CancelOnDrop type both binaries use (ADR-080 C2, #786)
Cross-turn KV / GDN prefix cache supported, text-only -- lattice serve's shared Metal worker calls generate_streaming_with_prefix_cache_and_cancel for every text Metal HTTP request, streaming and non-streaming alike; a vision-classified job routes to generate_multimodal_vision_with_cancel instead and never reaches the cache-aware call (crates/inference/src/serve/metal_worker.rs; see docs/cross-turn-cache.md). PR #619 wired chat_metal first; the HTTP surface followed supported -- PR #662 wired chat_completion_streaming_with_prefix_cache_and_cancel into the worker loop, reusing the previous turn's shared prefix instead of unconditionally reset_state()-ing (lattice_serve.rs:989-)
Grammar / JSON-schema constrained output not yet as an HTTP field -- response_format is parsed and rejected unless {"type":"text"} (reject_unsupported, lattice.rs:3203-); the underlying GenerateConfig.grammar: Option<Arc<GrammarEngine>> exists (ADR-046, qwen35_config.rs:669-673) but neither lattice chat nor lattice serve ever populates it from a request not yet -- issue #588. build_cfg hardcodes grammar: None (lattice_serve.rs); ChatReq now models response_format (#656) only to reject non-text values the same way lattice serve does -- "response_format.type '...' is not supported; use 'text'" -- not to implement it
LoRA adapters Multiple resident adapters via load/list/unload, selected by ordered lora: [{id, scale}] per chat request; absent or empty selects base. Metal only; CPU refuses by name. Same residency and selection contract, including streaming and non-streaming requests.
MTP / speculative decoding supported only via the LATTICE_MTP environment variable, not a CLI flag or request field -- GenerateConfig.enable_mtp: Option<bool> defaults to None ("defer to LATTICE_MTP env var", qwen35_config.rs:666-668,721) in both chat and serve. doctor reports has_mtp_tensors as an informational field only (lattice.rs:850,1098) supported only via the same LATTICE_MTP env-var default -- build_cfg also hardcodes enable_mtp: None (lattice_serve.rs:228); no per-request override in ChatReq
Reasoning budget / s1 (static forcing) supported (#831) -- a per-request reasoning_budget field is validated, normalized (positive values only -- Some(0) falls back to absent, but a malformed value, e.g. a non-numeric scalar, is strictly rejected with 400 invalid_request_body since both profiles honor the field -- see honoring_profile_still_rejects_malformed_sampling_fields), and applied to GenerateConfig.reasoning_budget via the shared serve::contract::normalize_request/ServeProfile::lattice (reasoning_budget_supported: true); no CLI flag or server-wide default, request-only. Previously accepted-and-silently-ignored -- see crates/inference/CHANGELOG.md ("Unreleased") supported (static) -- --reasoning-budget startup flag sets a server-wide default (lattice_serve.rs:737), and a per-request reasoning_budget field can override it, clamped to the 4096-token cache window (lattice_serve.rs:163-166,208-215)
Reasoning budget / s1 (adaptive, entropy-gated) not yet -- issue #500, tracked as extending #435 not yet -- issue #500
Embeddings supported two ways: CLI (embed binary, crates/embed/src/bin/embed.rs, a separate crate/binary entirely, text in / vector out, --json event mode, plus --image for a single inline image) and HTTP -- lattice serve exposes POST /v1/embeddings (text and/or inline-data-URI images, mixed batches, mean_visual/last_token pooling; see docs/serve-http-api.md) supported (#584) -- POST /v1/embeddings, text only, via a separately loaded --embedding-model BertModel (independent of the vision-language loader lattice serve uses); see the HTTP routes section above
Perplexity evaluation CLI-only, no serve equivalent -- eval_perplexity (CPU + Q4 + QuaRot dual-measurement modes, ADR-044 step 4) and ppl_metal (Metal-only, env-var driven: LATTICE_MODEL_DIR, PPL_TOKENS) not applicable -- this is an offline evaluation tool, not a request-shaped capability
Quantization (Q4 / QuaRot) workflow CLI-only, no serve equivalent -- quantize_q4 (stream BF16 -> Q4_0 .q4) and quantize_quarot (Hadamard-rotated Q4_0, ADR-044 step 3c-5). Both lattice chat/serve and lattice_serve consume a Q4 directory at load time (backend::detect_format, lattice.rs:100-138) but neither performs quantization itself CLI-only, no serve equivalent (same as previous column) -- quantization is an offline preprocessing step, not a serving-time capability in either binary
Metrics / observability not yet -- no /metrics route or equivalent in either subcommand of lattice.rs partial, not a real gap-filler for #583 -- lattice_serve prints a structured @@lattice {"ev":"http_request",...} line to stdout per request (emit_serve_event, lattice_serve.rs:608-628) for the Studio app's own parser; this is stdout event logging, not a queryable Prometheus/metrics HTTP endpoint. Issue #583 tracks the real /metrics route
/v1/models supported (ADR-080 C2, #786) -- GET /v1/models, single-entry list built from the shared lattice_inference::serve::models_list_body helper (lattice.rs, router()) supported -- GET /v1/models, single-entry list, same shared models_list_body helper (lattice_serve.rs)
GET / supported (ADR-080 C2, #786) -- engine-identity/endpoint-discovery document, shared lattice_inference::serve::root_body helper (lattice.rs, router()) supported -- same shared root_body helper (lattice_serve.rs)
GET /health supported -- always 200 {"status":"ok"} once listening, no dependency check (docs/serve-http-api.md) supported -- GET /health (lattice_serve.rs)
stop sequences supported -- string or array of up to 4 non-empty strings, parsed by the shared serve::contract::parse_stop_strings; validation order is documented in docs/serve-http-api.md supported (#831) -- ServeProfile::lattice_serve now also accepts and parses stop via the same shared parse_stop_strings, threading it into GenerateConfig.stop_strings alongside the fixed <|im_end|> token id (build_cfg, lattice_serve.rs:1652). Previously rejected outright with HTTP 400 "stop is not supported by this server"
seed supported -- req.seed passed straight through (lattice.rs:3319) supported -- req.seed passed straight through (lattice_serve.rs:225)
temperature supported -- validated to [0.0, 2.0] by the shared request contract, default 0.7 supported -- validated to [0.0, 2.0] by the same shared request contract, with the server-configured default
top_p supported -- validated to (0.0, 1.0] by the shared request contract, default 0.9 supported -- validated to (0.0, 1.0] by the same shared request contract, with the server-configured default
top_k accepted-and-ignored by ServeProfile::lattice (matches the pre-shared-contract DTO, which didn't model the field at all); generation keeps the fixed default of 50 supported -- mapped by ServeProfile::lattice_serve, with the server-wide default when omitted
repetition_penalty accepted-and-ignored by ServeProfile::lattice (matches the pre-shared-contract DTO, which didn't model the field at all); generation keeps the fixed default of 1.1 supported -- mapped by ServeProfile::lattice_serve, with the server-wide default when omitted
min_p engine-supported through SamplingConfig and canonical GenerateConfig; the HTTP request DTO does not expose it yet, so this surface uses the disabled default 0.0 same engine support and disabled server default; per-request HTTP wiring remains a later #587 slice
logprobs / top_logprobs supported, non-streaming only, CPU backend -- logprobs: true + top_logprobs (0-20), rejected if combined with stream: true (shared normalize_logprobs, crates/inference/src/serve/contract.rs); shipped in PR #620, closing #585. On the Metal/Q4 backend the request passes validation but fails with a generic 500 once generation starts: the shared cross-turn-cache-aware path's check_logprobs_not_set guard rejects it (see docs/serve-http-api.md "logprobs") not exposed -- ChatReq models the fields (#656) only to reject them explicitly: logprobs: true/top_logprobs gets HTTP 400 "logprobs/top_logprobs are not supported by this server" (shared reject_unsupported, crates/inference/src/serve/contract.rs) rather than being silently ignored; build_cfg still hardcodes logprobs: None for the requests that do proceed (lattice_serve.rs:878)
tools / tool_choice rejected with 400 if present, not silently dropped (shared reject_unsupported, crates/inference/src/serve/contract.rs) rejected with 400 if present (#656 closed the prior silent-ignore gap) -- ChatReq now models both fields explicitly so the shared reject_unsupported (crates/inference/src/serve/contract.rs) can name them: "tools and tool_choice are not supported by this server"
n (multiple completions) rejected with 400 if > 1 (shared reject_unsupported, crates/inference/src/serve/contract.rs) rejected with 400 if > 1 (#656) -- ChatReq.n is modeled explicitly and checked by the shared reject_unsupported (crates/inference/src/serve/contract.rs), matching lattice serve's contract
Non-text content parts (image/audio/file) one user-message image_url part is supported for non-streaming requests when the loaded Metal checkpoint passes exact vision-config/token/inventory preflight; PNG/JPEG data: URIs only, 48,000 decoded bytes maximum, remote URLs rejected. Text-only/CPU models return 400 vision_unsupported; audio/file remain unsupported. identical shared-contract behavior: vision-capable Metal checkpoints accept the same bounded, non-streaming inline PNG/JPEG shape, while text-only models return the same 400 code; audio/file remain unsupported.
Unknown/malformed role rejected with 400, both for named-but-unsupported roles ("tool", "developer") and any other unrecognized role; production runs through shared serve::contract::normalize_request, which produces contract-owned NormalizedChatMessage values consumed by the single serve::into_engine_chat_messages adapter (see docs/serve-http-api.md "Request fields") rejected with the same shared contract and error codes (#641 closed by PR #656); parse_chat_req performs bounded structural JSON parsing, then normalize_request validates roles into NormalizedChatMessage before the single contract-to-engine adapter runs; the remaining local MessageRole is #[cfg(test)]-only historical error-contract coverage, not a production conversion path
model field validation rejected with 400 if it doesn't match the served model id (lattice.rs:3291-) rejected with 400 model_not_found if it doesn't match the served model id (lattice_serve.rs, cm_lattice_serve_model_mismatch_rejected)
Max context / token-budget cap max_tokens_cap hardcoded to 4096 in lattice serve's main() (not a CLI flag); Metal/Q4 backend additionally caps total context at MetalChatBackend::MAX_CACHE_LEN (4096) regardless of the model's configured max (docs/serve-http-api.md) supported, derived from model config (#551, closed) -- AppState.model_max_context is set at load time from the model's max_position_embeddings (model_context_from_config, lattice_serve.rs:828-), not a hard-coded constant; FALLBACK_MODEL_MAX_CONTEXT (4096, lattice_serve.rs:178) is used only when no config is available
Auth / rate limiting not implemented -- no Authorization handling, no request-count or concurrency-limiting middleware (docs/serve-http-api.md) not implemented -- same absence, no auth/rate-limit layer in the router shown above
Admission control (Metal pending-job cap) supported -- both binaries share the same MetalWorkerClient/DEFAULT_MAX_PENDING_JOBS = 32 default (crates/inference/src/serve/metal_worker.rs, issue #932), overridable per-process via --max-pending (see docs/serve-http-api.md), rejecting with HTTP 503 server_busy at the worker's admission boundary before any Metal generation work happens -- the HTTP handler has already tokenized the rendered prompt during request preparation ahead of that point, so this is not an end-to-end pre-tokenization guarantee; applies to both streaming and non-streaming requests, checked before an SSE stream is committed to on the streaming path. Not a rate limiter (see "Auth / rate limiting" above) -- bounds outstanding work on the one shared worker thread, not requests per client or unit time. No cap on the CPU backend supported -- same shared MetalWorkerClient admission cap (issue #932), same --max-pending override

Summary: where the two serve surfaces diverge today

Capabilities present in lattice serve (the CLI's HTTP subcommand) but missing from lattice_serve: logprobs/top_logprobs is still not implemented on lattice_serve (the field is now explicitly rejected with 400 rather than silently ignored -- #656 -- but it doesn't run). As of #831, stop sequences are no longer in this list -- ServeProfile::lattice_serve now accepts and parses stop too (see the matrix row above). /v1/embeddings is no longer in this list either: both surfaces now register the route, but the two implementations are independent -- lattice serve pools text and inline-data-URI images through a loaded vision-language checkpoint, while lattice_serve (as of #584) embeds text only, through a separately loaded BertModel (see the matrix row above and the HTTP routes section for both).

Capabilities present in lattice_serve but missing from lattice serve: per-request top_k/repetition_penalty control. As of PR #619, cross-turn/prefix KV cache reuse over HTTP is no longer in this list -- lattice serve's Metal worker now calls the cache-aware generation path too (see the matrix row above). As of #831, static reasoning-budget forcing (reasoning_budget request field) is no longer in this list -- ServeProfile::lattice now accepts and applies it too (see the matrix row above). As of ADR-080 C2 (#786), /v1/models, GET /, and disconnect-cancellation on the streaming path are no longer in this list -- lattice serve now has all three, built on the same shared lattice_inference::serve contract lattice_serve uses.

The shared contract also owns the Metal worker lifecycle (#833). In both binaries, each client's Drop explicitly closes its job sender before automatic field destruction can release its owner clone. The final owner transfers the worker's join handle to a detached reaper and waits up to two seconds for its result. That deadline covers the full join, including thread-local destructors; if the worker or a destructor is still running when it expires, the reaper remains detached so process shutdown cannot hang indefinitely. Both binaries use the same server runner, which gives tracked HTTP/1 connections up to five seconds to drain on SIGINT and Unix SIGTERM. After any drain timeout it aborts remaining connection tasks, allows up to three seconds for cancellation cleanup, and then exits with status 1 even if cleanup completed. That hard-exit fallback can truncate in-flight responses, leave files partially written, and discard unflushed telemetry because it skips Rust destructors.

As of #656, tools/tool_choice/n/non-text response_format/unknown-role/non-text-content-part rejection now behave identically on both HTTP surfaces (all 400, none silently ignored or coerced) -- this was previously the biggest divergence between the two surfaces and it has since closed.

Studio request compatibility

lattice_serve now returns HTTP 400 invalid_temperature for temperature values below 0.0 or above 2.0, and HTTP 400 invalid_top_p for top_p values at or below 0.0 or above 1.0. These requests previously passed daemon validation and could reach generation. A top-level field outside the shared request DTO (e.g. presence_penalty, frequency_penalty, logit_bias, user) is ignored rather than rejected, matching standard OpenAI-compatible server behavior. The daemon's stop, model-name, max-token clamping, top_k, repetition_penalty, reasoning-budget, structured-output, and logprobs policies are unchanged.

Capabilities missing from both: grammar/JSON-schema constrained output over HTTP (#588), a real /metrics endpoint (#583), adaptive/entropy-gated reasoning budget (#500), per-request min_p wiring (engine support is the first #587 slice). LoRA residency and per-request selection are available over HTTP on Metal.

Gaps tracked for this document: #641 (closed by PR #656) -- lattice_serve silently coerced unknown roles to "user" and dropped non-text content parts, where lattice serve rejected both with HTTP 400.

Fixture manifest (#654)

Every row in the Capability matrix above must map to a #[test] fn (a fixture ID, cited by name) in the cited surface's own binary (or, for behavior shared by both binaries through the lattice_inference::serve::metal_worker owner, that shared module -- #832), or carry an explicit NOT-TESTABLE-YET justification. scripts/check-capability-matrix.sh (wired into make lint-docs) greps crates/inference/src/bin/{lattice,lattice_serve,chat_metal}.rs and crates/inference/src/serve/metal_worker.rs for fn <fixture id> and fails the build if a cited ID no longer exists, or if a row has no fixture cell at all -- so a row that goes stale (a fixture renamed/removed without updating this table) fails closed instead of silently drifting.

Fixtures test the request-validation/acceptance contract each row describes (a 400 with a given code, a value resolving through to GenerateConfig, a request reaching the job queue) without a loaded model or live GPU -- both binaries' handlers run every check in this table before they ever touch the model (see validate_chat_request in lattice.rs, and the Job-queue split in lattice_serve.rs). Actual token generation, real Metal dispatch, and CLI-tool output content are out of scope for this harness and marked NOT-TESTABLE-YET; those are covered instead by the e2e parity gate (docs/../.github/workflows/e2e-parity.yml) and manual lattice doctor/ppl_metal runs.

Capability row lattice serve fixture (lattice.rs) lattice_serve fixture (lattice_serve.rs)
Chat completion (single/multi-turn) cm_serve_model_match_passes_model_check; request-shape guards chat_completions_rejects_message_flood (MAX_MESSAGE_COUNT bound), chat_completions_matches_shared_parity_table; config-capture chat_completions_non_streaming_observation_captures_real_config_and_prompt, chat_completions_non_streaming_observation_captures_real_stopped_false cm_lattice_serve_model_mismatch_rejected; message-content parsing message_content_plain_string_normalizes, message_content_parts_concatenate_in_order; chat_completions_matches_shared_parity_table; config-capture chat_completions_non_streaming_observation_captures_real_config_and_messages, chat_completions_non_streaming_observation_captures_real_stopped_false; finish_reason mapping chat_completions_non_streaming_finish_reason_stop_when_engine_stopped, chat_completions_non_streaming_finish_reason_length_when_not_stopped, chat_completions_streaming_finish_reason_stop_when_engine_stopped, chat_completions_streaming_finish_reason_length_when_not_stopped
Streaming (SSE) NOT-TESTABLE-YET: stream: true only changes which response arm the handler takes; the SSE body itself needs a running model to produce tokens (reject_unsupported_stream_true_ok/reject_unsupported_stream_false_ok cover the field's acceptance, not the stream). Mid-stream failure envelope IS fixtured via a fake-generate seam: chat_completions_streaming_failure_emits_error_event; shared metal_worker classification underlying it: generation_failure_is_reported_as_failed_not_complete NOT-TESTABLE-YET: same reason; Phase/SSE framing needs live worker output. Mid-stream failure envelope: chat_completions_streaming_failure_emits_error_event, chat_completions_streaming_failure_records_failed_metric
Disconnect cancellation (mid-stream) chat_completions_streaming_disconnect_stops_generation (HTTP-level, test-utils feature: drives the real chat_completions -> body_stream -> generate_streaming_with_cancel composition against a tiny CPU model via BodyExt::frame); generate_streaming_with_cancel_true_before_prefill_returns_interrupt, generate_streaming_with_cancel_true_after_prefill_returns_interrupt, generate_streaming_with_cancel_mid_decode_stops_early_fast_path, generate_streaming_with_cancel_mid_decode_stops_early_stop_string_path (engine-level, model/qwen35/generation.rs, one fixture per checkpoint); chat_completions_streaming_disconnect_cancellation_reaches_generator_post_drop (proves the cancel signal reaches the generator closure itself, not just the HTTP layer) queued_job_cancelled_before_dequeue_sends_exactly_one_cancelled_event, running_job_cancelled_midstream_stops_early_and_worker_survives, running_job_cancelled_during_prefill_like_phase_never_calls_on_token (issue #832: moved to the shared lattice_inference::serve::metal_worker owner both lattice_serve.rs and, once its own migration lands on this branch, lattice.rs's Metal backend route through)
Cross-turn KV / GDN prefix cache NOT-TESTABLE-YET: requires a loaded model + live Metal GPU; tracked by #462 NOT-TESTABLE-YET: same -- the cache-aware call path (lattice_serve.rs:989-) needs a live MetalQwen35State; covered by manual/e2e verification, not this harness
Grammar / JSON-schema constrained output reject_unsupported_response_format_json, reject_unsupported_response_format_text_ok chat_completions_json_response_format_400; structured-output validation chat_completions_rejects_duplicate_root_type_400, chat_completions_rejects_duplicate_nested_properties_400, chat_completions_rejects_duplicate_key_inside_property_schema_400, chat_completions_structured_missing_strict_400, chat_completions_structured_strict_false_400, chat_completions_structured_streaming_400, chat_completions_structured_pattern_keyword_400, chat_completions_structured_enum_keyword_400, chat_completions_structured_applies_grammar_and_marker, chat_completions_structured_blocked_constraint_500, chat_completions_structured_failed_with_blocked_wording_stays_internal_error, chat_completions_structured_length_limit_500, chat_completions_structured_validation_failed_500
LoRA adapters a_cpu_backend_refuses_an_adapter_by_name, lora_unload_on_a_cpu_backend_refuses_by_name, lora_list_on_cpu_refuses_by_name, lora_selection_on_cpu_refuses_by_name, lora_duplicate_id_is_400_before_unified_admission. Shared registry: residents_are_listed_without_applying_and_ids_are_never_reused, same_mixture_blends_once_changed_order_or_scale_reblends, absent_selection_restores_base_output_with_residents_present. lora_unknown_id_is_400_before_streaming, lora_duplicate_id_is_400_before_admission, lora_list_reads_confirmed_index, lora_selection_reaches_worker, lora_load_nonexistent_file_400, lora_unload_without_a_worker_is_503.
MTP / speculative decoding NOT-TESTABLE-YET: enable_mtp is env-var-only, not a request field; nothing HTTP-shaped to fixture NOT-TESTABLE-YET: same
Reasoning budget / s1 (static forcing) both_profiles_apply_reasoning_budget, context_window_accounts_for_reasoning_budget (shared serve::contract unit tests), chat_completions_non_streaming_observation_captures_real_reasoning_budget (config-capture, proves the field reaches the real GenerateConfig) NOT-TESTABLE-YET: reasoning_budget clamping needs build_cfg end-to-end against a real model_max_context; only build_cfg_clamps_to_runtime_context exercises the clamp math, not the full request path
Reasoning budget / s1 (adaptive, entropy-gated) NOT-TESTABLE-YET: issue #500, not implemented on either surface NOT-TESTABLE-YET: issue #500
Embeddings router-level coverage exists (embeddings_route module, lattice.rs's serve bin, #[cfg(feature = "test-utils")]): no_embedding_model_loaded_fails_closed, happy_path_text_only, happy_path_image, mixed_batch_preserves_input_order; these names fall outside this script's fixture-ID prefix convention (see the prefix list at the top of scripts/check-capability-matrix.sh), so they are not mechanically cross-checked by it, but the tests are real and pass in CI. The separate embed CLI binary/crate (text/image in, vector out) is out of this harness's scope entirely partially fixtured (#584): request parsing/validation/response-building is pure and metal-gpu-independent, covered by unit tests in lattice_inference::serve::embeddings (none of the fixture-ID prefixes this script tracks apply to that module's names, so it is not cited by prefix below); the HTTP handler, AppState wiring, and tests in lattice_serve.rs (embeddings_missing_input_400, embeddings_no_model_loaded_503, and others -- see the module itself) live in this binary's metal-gpu-gated imp module like the rest of its tests, so NOT-TESTABLE-YET in that harness sense, same as every other row gated the same way
Perplexity evaluation NOT-TESTABLE-YET: offline CLI tool (eval_perplexity/ppl_metal), not a request-shaped capability not applicable
Quantization (Q4/QuaRot) workflow NOT-TESTABLE-YET: offline CLI tool, not a serving-time capability not applicable
Metrics / observability NOT-TESTABLE-YET: no /metrics route on either surface (issue #583) NOT-TESTABLE-YET: emit_serve_event's stdout format is a Studio-app-internal wire contract, not yet fixtured here
/v1/models NOT-TESTABLE-YET: list_models returns a static single-entry list from AppState.model_id via the shared models_list_body helper; route now exists (ADR-080 C2, #786) but not yet fixtured at the router level beyond the shared root_body_shape/models_list_body_shape unit tests in serve/mod.rs NOT-TESTABLE-YET: list_models returns a static single-entry list from AppState.model_id; trivial but not yet fixtured, tracked as a follow-up rather than blocking #654
GET /health not applicable -- trivial liveness route, no request-shaped contract to fixture not applicable, same reason
stop sequences parse_stop_strings_null_gives_empty, parse_stop_strings_single_string_gives_vec_of_one, parse_stop_strings_array_of_two_accepted, parse_stop_strings_empty_array_rejected, parse_stop_strings_array_over_four_rejected, parse_stop_strings_array_with_number_rejected, parse_stop_strings_empty_string_element_rejected, parse_stop_strings_empty_string_scalar_rejected, parse_stop_strings_array_exactly_four_accepted (9 fixtures), reject_unsupported_stop_now_accepted, cm_serve_stop_sequences_accepted_end_to_end chat_completions_stop_is_accepted_and_reaches_generate_config (real-worker, config-capture)
seed NOT-TESTABLE-YET: pass-through with no validation logic to fixture NOT-TESTABLE-YET: same
temperature shared parity rows temperature_boundary_zero_accepted, temperature_boundary_two_accepted, temperature_out_of_range_rejected; shared contract unit tests; validate_temperature_rejects_negative, validate_temperature_rejects_above_two, validate_temperature_accepts_boundary, validate_temperature_none_uses_default same shared parity rows through the daemon router
top_p shared parity rows top_p_boundary_one_accepted, top_p_zero_rejected, top_p_above_one_rejected; shared contract unit tests; validate_top_p_rejects_zero, validate_top_p_rejects_above_one, validate_top_p_accepts_one, validate_top_p_none_uses_default same shared parity rows through the daemon router
top_k shared profile unit tests (profiles_preserve_sampling_extension_policy, lattice_profile_tolerates_malformed_ignored_sampling_fields -- accepted-and-ignored, not rejected) build_cfg_clamps_to_runtime_context plus shared profile unit tests (honoring_profile_still_rejects_malformed_sampling_fields)
repetition_penalty shared profile unit tests (profiles_preserve_sampling_extension_policy, lattice_profile_tolerates_malformed_ignored_sampling_fields -- accepted-and-ignored, not rejected) shared profile unit tests (honoring_profile_still_rejects_malformed_sampling_fields)
min_p engine tests min_p_filters_by_probability_relative_to_maximum, min_p_runs_before_top_p_and_renormalizes_survivors, and min_p_is_active_and_identical_across_cpu_sampling_paths; no HTTP request field yet same engine tests plus compact_metal_candidate_path_applies_min_p; no HTTP request field yet
logprobs / top_logprobs validate_logprobs_absent_disables_capture, validate_logprobs_false_disables_capture, validate_logprobs_true_no_top_logprobs_defaults_to_zero, validate_logprobs_true_with_top_logprobs_ok, validate_logprobs_top_logprobs_at_boundary_twenty_ok, validate_logprobs_top_logprobs_over_twenty_rejected, validate_logprobs_top_logprobs_without_logprobs_true_rejected, validate_logprobs_top_logprobs_with_logprobs_false_rejected (8 fixtures), cm_serve_logprobs_resolved_end_to_end; unsupported-combo rejection reject_unsupported_stream_and_logprobs_rejected, reject_unsupported_logprobs_true_ok, reject_unsupported_logprobs_false_ok chat_completions_logprobs_400
tools / tool_choice reject_unsupported_tools_rejected, reject_unsupported_tool_choice_rejected chat_completions_tools_400, chat_completions_tool_choice_400
n (multiple completions) reject_unsupported_n_gt_1, reject_unsupported_n_1_ok chat_completions_n_greater_than_one_400
Non-text content parts (image/audio/file) vision_content_parts_1135::{vision_model_accepts_image_and_enqueues_it_on_the_shared_worker,text_only_model_rejects_the_same_image_with_capability_code}, shared contract/adapter/worker tests chat_completions_vision_model_enqueues_image_on_shared_worker, chat_completions_image_url_400, chat_completions_oversized_part_400, shared contract/adapter/worker tests; model-gated production route: vision_serve_e2e_test; content-part rejection message_content_image_url_rejected, message_content_unknown_part_rejected
Unknown/malformed role to_chat_messages_rejects_invalid_role, to_chat_messages_rejects_tool_role, to_chat_messages_rejects_developer_role chat_completions_unknown_role_400, message_role_unknown_rejected, message_role_tool_and_developer_rejected_as_unsupported_feature, chat_completions_tool_and_developer_role_400_unsupported_feature
model field validation cm_serve_model_mismatch_rejected, cm_serve_model_match_passes_model_check cm_lattice_serve_model_mismatch_rejected
Max context / token-budget cap validate_max_tokens_rejects_zero, validate_max_tokens_rejects_above_cap, validate_max_tokens_alias_conflict_rejected, validate_max_tokens_at_exactly_cap_ok, validate_max_tokens_uses_default_when_absent, validate_max_tokens_alias_agrees, validate_max_tokens_max_completion_only_ok; context_window_accepts_boundary_and_rejects_overflow, context_window_accounts_for_reasoning_budget (#831: shared saturating full-window formula, prompt + max_new_tokens + reasoning_budget + 1 <= max_context, used by both binaries' preflight); chat_completions_streaming_context_overflow_returns_400_before_committing_sse build_cfg_clamps_to_runtime_context, chat_completions_conflicting_max_tokens_400, model_context_uses_config_max_position_embeddings, model_context_falls_back_to_4096_when_absent, check_prompt_fits_window_rejects_when_prompt_plus_decode_overflows, check_prompt_fits_window_accepts_ordinary_prompt_unclamped (2 fixtures), chat_completions_max_tokens_zero_400, chat_completions_streaming_context_overflow_returns_400_before_committing_sse, chat_completions_streaming_context_overflow_matches_lattice_real_router, build_cfg_aliases_max_completion_tokens_when_max_tokens_absent
Auth / rate limiting not applicable -- no auth/rate-limit layer exists on either surface to fixture not applicable, same
Admission control (Metal pending-job cap) chat_completions_streaming_returns_503_before_sse_commit_at_cap, chat_completions_non_streaming_returns_503_server_busy_not_500_at_cap chat_completions_returns_503_json_envelope_when_admission_cap_reached
Request precedence / order cm_serve_unsupported_feature_rejected_before_model_check, cm_serve_empty_messages_rejected, cm_serve_last_message_not_user_rejected (previously untested: these three checks ran only inline in chat_completions with no dedicated fixture at all); cm_serve_context_window_checked_before_stop_parsing pins that the context-window preflight runs before stop-sequence parsing, restoring the original inline order across the prepare_chat_request extraction shared normalize_request runs unsupported-feature checks before model/message/sampling validation; parse_chat_req only performs bounded structural JSON parsing -- bounds enforced by parse_chat_req_rejects_too_many_parts_before_typed_parse, parse_chat_req_accepts_max_parts_boundary, parse_chat_req_rejects_oversized_text_part_before_typed_parse. The chat_completions_*_400 fixtures above cover this ordering, and the shared context-precedence fixture pins context validation before stop parsing

chat_metal --json --serve fixtures

chat_metal is not part of the two-column OpenAI-shaped matrix above -- its --json --serve mode speaks a simpler newline-delimited-JSON stdin protocol ({"prompt":...} in, @@lattice {"ev":"gen_token",...} out), not the /v1/chat/completions wire shape, and it has no stop/messages/tools/role fields at all: the client resends the full ChatML-rendered history as a single prompt string every request (see the module doc comment at the top of chat_metal.rs). Its request-parsing was previously inline in the serve loop with no tests at all; #654 factored it into parse_serve_request_line (chat_metal.rs) specifically so it could be fixtured:

Wire behavior Fixture
prompt-only request falls back to all CLI defaults cm_chat_metal_serve_prompt_only_uses_all_defaults
Per-request field overrides (max_tokens/temperature/top_k/top_p/repetition_penalty/seed/reasoning_budget) cm_chat_metal_serve_per_request_overrides_all_fields
Malformed JSON line cm_chat_metal_serve_malformed_json_rejected
Missing prompt field cm_chat_metal_serve_missing_prompt_rejected
reasoning_budget: 0 treated as absent (not an explicit zero-budget override) cm_chat_metal_serve_zero_reasoning_budget_falls_back_to_default
No stop/history fields exist on the wire shape -- an extra stop key is inert, not an error cm_chat_metal_serve_stateless_no_stop_strings_or_history_fields
Actual Metal generation, @@lattice gen_token event stream content, cross-turn cache reuse behavior NOT-TESTABLE-YET: requires a loaded model + live Metal GPU

See also

  • docs/serve-http-api.md -- full lattice serve HTTP API reference (request/response shapes, validation order, error envelopes, streaming details)
  • docs/cross-turn-cache.md -- cross-turn KV/GDN prefix cache: what calls it and what doesn't
  • docs/q4-quantization.md -- Q4/QuaRot quantization workflow and lattice doctor output
  • #604 -- usage-example/cookbook coverage tracking, including #601 (lattice_serve HTTP API walkthrough) and #602 (CLI tool walkthroughs)