This document tracks which capabilities exist in each of lattice's two OpenAI-shaped binaries, so new features don't silently land in only one of them. It is the reference #589 asked for.
Same trap as docs/serve-http-api.md and docs/cross-turn-cache.md describe -- read this first:
latticeCLI (crates/inference/src/bin/lattice/--main.rsdispatch,chat.rsREPL,doctor.rspreflight,serve.rsHTTP API; split from a singlelattice.rsfile in #834) -- one binary, three subcommands:lattice chat(interactive REPL),lattice serve(an HTTP server, OpenAI-compatible),lattice doctor(preflight check, no model weights loaded). This is the README-documented, actively developed surface. Itsservesubcommand is itself an HTTP server -- when a row below says "supported" for this column via HTTP, it meanslattice serve's router, notlattice chat. Below, unqualifiedlattice.rs:NNNline citations predate the #834 split and are stale on file/line but not on subcommand/behavior; re-derive against the split files rather than trusting the number.lattice_serve(crates/inference/src/bin/lattice_serve.rs) -- a separate, standalone binary purpose-built as the HTTP daemon the macOS Lattice Studio app spawns and talks to (introduced PR #435). Metal-GPU-only, its own fixed route list, its own request struct.
Every "supported" claim below cites a file:line. Every "not yet" cites the tracking issue.
As of ADR-080 cluster C2 (#786) -- see crates/inference/src/serve/mod.rs
for the shared contract (error envelope, finish_reason mapping, max_tokens == 0 rejection,
/+/v1/models response bodies, disconnect-cancellation guard) both binaries now build their
routers on top of.
Every row below is backed by an executable fixture or an explicit
not-testable-yet justification: see the Fixture manifest
section, checked by scripts/check-capability-matrix.sh (wired into make lint-docs). A row that names a fixture ID citing a test function that no
longer exists fails that check closed (#654).
| Subcommand | Flags | Evidence |
|---|---|---|
chat |
--model (required), --max-tokens (default 256), --temperature (default 0.7), --tokenizer-dir |
lattice.rs:22-39 |
serve |
--model (required), --host (default 127.0.0.1), --port (default 8080), --max-tokens (default 256, per-request default), --model-id, --tokenizer-dir |
lattice.rs:40-63 |
doctor |
--model (required), --context (optional target length), --tokenizer-dir |
lattice.rs:64-80 |
No --lora, --reasoning-budget, --top-k, or --repetition-penalty flags exist on any
subcommand -- chat only ever builds a GenerateConfig::default() and then assigns
max_new_tokens and temperature, nothing else (bin/lattice/chat.rs:149-151), so
top_k/repetition_penalty/seed/grammar/MTP/
reasoning-budget are all fixed at their Default values (qwen35_config.rs:2429-2447: top_k: 50, top_p: 0.9, repetition_penalty: 1.1, enable_mtp: None, grammar: None, reasoning_budget: None)
for the whole interactive session -- there is no way to change them without editing source.
Complete router (lattice.rs, mod serve, router()):
Router::new()
.route("/", get(root))
.route("/health", get(health))
.route("/v1/models", get(list_models))
.route("/v1/chat/completions", post(chat_completions))
.route("/v1/embeddings", post(embeddings))
.route("/v1/lora", get(lora_list))
.route("/v1/lora/load", post(lora_load))
.route("/v1/lora/unload", post(lora_unload))
.layer(DefaultBodyLimit::max(REQUEST_BODY_LIMIT_BYTES))Both GET / and GET /v1/models landed in ADR-080 C2 (#786):
lattice serve previously had neither route. GET / and /v1/models' response bodies are now
built from the shared lattice_inference::serve::root_body/models_list_body helpers so both
binaries return byte-identical shapes for the same inputs. POST /v1/embeddings is always
installed on the router, but every request 400s with vision_unsupported unless the server was
started with --model pointed at a vision-language checkpoint (see
docs/serve-http-api.md). No /v1/completions, no admin/metrics route.
Complete router (lattice_serve.rs, router(), extracted from run() so it's
directly testable via tower::ServiceExt::oneshot):
Router::new()
.route("/", get(root))
.route("/health", get(health))
.route("/v1/models", get(list_models))
.route("/v1/chat/completions", post(chat_completions))
.route("/v1/embeddings", post(embeddings))
.route("/v1/lora", get(lora_list))
.route("/v1/lora/load", post(lora_load))
.route("/v1/lora/unload", post(lora_unload))GET /v1/models returns a single-entry OpenAI model list for the one loaded model, via the same
shared models_list_body helper lattice serve now uses. Deliberately does NOT install
DefaultBodyLimit: this binary enforces the same REQUEST_BODY_LIMIT_BYTES cap manually inside
chat_completions via to_bytes, a documented intentional divergence in mechanism (same limit
value, different enforcement path) -- see the cross-binary parity table in serve/mod.rs
(CHAT_COMPLETIONS_PARITY_CASES) for the full list of aligned-vs-documented-divergent outcomes.
POST /v1/embeddings (issue #584) is always registered, but only answers requests when the
binary was started with --embedding-model <dir> (or LATTICE_SERVE_EMBEDDING_MODEL): without
it, a request that otherwise passes content-type, body-size, JSON, and input validation gets
HTTP 503 embedding_model_not_loaded rather than the route not existing at all, so its absence
is a discoverable server state rather than a bare 404. Content-type/body-size/parse/field
validation run first and take precedence -- a request with a non-JSON Content-Type, an
oversized body, malformed JSON, or an invalid/unsupported field still gets its own 415, 413, or
400 even with no embedding model loaded. The route computes
embeddings via lattice_inference::model::bert::BertModel directly (CPU-only), not through
lattice-embed's EmbeddingService trait -- lattice-embed already depends on
lattice-inference for BertModel/QwenModel, so a reverse edge from this crate back onto
lattice-embed is a cyclic package dependency Cargo refuses to build (confirmed with
cargo tree -p lattice-inference after adding the edge experimentally: "cyclic package
dependency ... lattice-embed ... depends on itself"). An optional --embedding-pooling mean|cls flag selects the pooling strategy (mean is BertModel's own default; BGE-family
checkpoints need cls). Request/response DTOs, input normalization (single string or array,
untagged), and response building are in the shared, GPU-independent
lattice_inference::serve::embeddings module so that logic is unit-testable without a loaded
model or the metal-gpu feature; the HTTP handler, CLI flags, and AppState wiring live in
lattice_serve.rs itself, gated the same as the rest of this binary's imp module (macOS +
metal-gpu feature).
| Capability | lattice CLI |
lattice_serve HTTP |
|---|---|---|
| Chat completion (single/multi-turn) | supported -- lattice chat (REPL, lattice.rs:2206-2352) and lattice serve POST /v1/chat/completions (chat_completions/chat_completions_with_request, lattice.rs:3616-) |
supported -- POST /v1/chat/completions (lattice_serve.rs:425-) |
| Streaming (SSE) | supported -- lattice serve only, stream: true (lattice.rs, see docs/serve-http-api.md "streaming"); lattice chat's REPL has no streaming concept (prints the full response once) |
supported -- stream: true, its own SSE phase machine (lattice_serve.rs:415-) |
| Disconnect cancellation (mid-stream) | supported, both backends -- ADR-080 C2 (#744, #786) wired the shared lattice_inference::serve::CancelOnDrop/cancel_pair guard into the CPU streaming path: cancel_guard is captured by body_stream's flat_map closure so it drops (flipping cancel_rx) the instant axum drops the SSE response on client disconnect, and generate_streaming_with_cancel polls it independently of on_token at four checkpoints (pre-prefill, post-prefill, and the top of every decode-loop iteration, both the fast and stop-string paths). The Metal/Q4 path already had this via the same shared guard threaded through cancel_rx. Previously the CPU path had none (its delta closure ignored the send result, so generation ran to max_tokens after a disconnect) -- that gap is what #744 closed |
supported -- CancelOnDrop/watch::channel (PR #552/#606), now built on the same shared lattice_inference::serve::CancelOnDrop type both binaries use (ADR-080 C2, #786) |
| Cross-turn KV / GDN prefix cache | supported, text-only -- lattice serve's shared Metal worker calls generate_streaming_with_prefix_cache_and_cancel for every text Metal HTTP request, streaming and non-streaming alike; a vision-classified job routes to generate_multimodal_vision_with_cancel instead and never reaches the cache-aware call (crates/inference/src/serve/metal_worker.rs; see docs/cross-turn-cache.md). PR #619 wired chat_metal first; the HTTP surface followed |
supported -- PR #662 wired chat_completion_streaming_with_prefix_cache_and_cancel into the worker loop, reusing the previous turn's shared prefix instead of unconditionally reset_state()-ing (lattice_serve.rs:989-) |
| Grammar / JSON-schema constrained output | not yet as an HTTP field -- response_format is parsed and rejected unless {"type":"text"} (reject_unsupported, lattice.rs:3203-); the underlying GenerateConfig.grammar: Option<Arc<GrammarEngine>> exists (ADR-046, qwen35_config.rs:669-673) but neither lattice chat nor lattice serve ever populates it from a request |
not yet -- issue #588. build_cfg hardcodes grammar: None (lattice_serve.rs); ChatReq now models response_format (#656) only to reject non-text values the same way lattice serve does -- "response_format.type '...' is not supported; use 'text'" -- not to implement it |
| LoRA adapters | Multiple resident adapters via load/list/unload, selected by ordered lora: [{id, scale}] per chat request; absent or empty selects base. Metal only; CPU refuses by name. |
Same residency and selection contract, including streaming and non-streaming requests. |
| MTP / speculative decoding | supported only via the LATTICE_MTP environment variable, not a CLI flag or request field -- GenerateConfig.enable_mtp: Option<bool> defaults to None ("defer to LATTICE_MTP env var", qwen35_config.rs:666-668,721) in both chat and serve. doctor reports has_mtp_tensors as an informational field only (lattice.rs:850,1098) |
supported only via the same LATTICE_MTP env-var default -- build_cfg also hardcodes enable_mtp: None (lattice_serve.rs:228); no per-request override in ChatReq |
| Reasoning budget / s1 (static forcing) | supported (#831) -- a per-request reasoning_budget field is validated, normalized (positive values only -- Some(0) falls back to absent, but a malformed value, e.g. a non-numeric scalar, is strictly rejected with 400 invalid_request_body since both profiles honor the field -- see honoring_profile_still_rejects_malformed_sampling_fields), and applied to GenerateConfig.reasoning_budget via the shared serve::contract::normalize_request/ServeProfile::lattice (reasoning_budget_supported: true); no CLI flag or server-wide default, request-only. Previously accepted-and-silently-ignored -- see crates/inference/CHANGELOG.md ("Unreleased") |
supported (static) -- --reasoning-budget startup flag sets a server-wide default (lattice_serve.rs:737), and a per-request reasoning_budget field can override it, clamped to the 4096-token cache window (lattice_serve.rs:163-166,208-215) |
| Reasoning budget / s1 (adaptive, entropy-gated) | not yet -- issue #500, tracked as extending #435 | not yet -- issue #500 |
| Embeddings | supported two ways: CLI (embed binary, crates/embed/src/bin/embed.rs, a separate crate/binary entirely, text in / vector out, --json event mode, plus --image for a single inline image) and HTTP -- lattice serve exposes POST /v1/embeddings (text and/or inline-data-URI images, mixed batches, mean_visual/last_token pooling; see docs/serve-http-api.md) |
supported (#584) -- POST /v1/embeddings, text only, via a separately loaded --embedding-model BertModel (independent of the vision-language loader lattice serve uses); see the HTTP routes section above |
| Perplexity evaluation | CLI-only, no serve equivalent -- eval_perplexity (CPU + Q4 + QuaRot dual-measurement modes, ADR-044 step 4) and ppl_metal (Metal-only, env-var driven: LATTICE_MODEL_DIR, PPL_TOKENS) |
not applicable -- this is an offline evaluation tool, not a request-shaped capability |
| Quantization (Q4 / QuaRot) workflow | CLI-only, no serve equivalent -- quantize_q4 (stream BF16 -> Q4_0 .q4) and quantize_quarot (Hadamard-rotated Q4_0, ADR-044 step 3c-5). Both lattice chat/serve and lattice_serve consume a Q4 directory at load time (backend::detect_format, lattice.rs:100-138) but neither performs quantization itself |
CLI-only, no serve equivalent (same as previous column) -- quantization is an offline preprocessing step, not a serving-time capability in either binary |
| Metrics / observability | not yet -- no /metrics route or equivalent in either subcommand of lattice.rs |
partial, not a real gap-filler for #583 -- lattice_serve prints a structured @@lattice {"ev":"http_request",...} line to stdout per request (emit_serve_event, lattice_serve.rs:608-628) for the Studio app's own parser; this is stdout event logging, not a queryable Prometheus/metrics HTTP endpoint. Issue #583 tracks the real /metrics route |
/v1/models |
supported (ADR-080 C2, #786) -- GET /v1/models, single-entry list built from the shared lattice_inference::serve::models_list_body helper (lattice.rs, router()) |
supported -- GET /v1/models, single-entry list, same shared models_list_body helper (lattice_serve.rs) |
GET / |
supported (ADR-080 C2, #786) -- engine-identity/endpoint-discovery document, shared lattice_inference::serve::root_body helper (lattice.rs, router()) |
supported -- same shared root_body helper (lattice_serve.rs) |
GET /health |
supported -- always 200 {"status":"ok"} once listening, no dependency check (docs/serve-http-api.md) |
supported -- GET /health (lattice_serve.rs) |
stop sequences |
supported -- string or array of up to 4 non-empty strings, parsed by the shared serve::contract::parse_stop_strings; validation order is documented in docs/serve-http-api.md |
supported (#831) -- ServeProfile::lattice_serve now also accepts and parses stop via the same shared parse_stop_strings, threading it into GenerateConfig.stop_strings alongside the fixed <|im_end|> token id (build_cfg, lattice_serve.rs:1652). Previously rejected outright with HTTP 400 "stop is not supported by this server" |
seed |
supported -- req.seed passed straight through (lattice.rs:3319) |
supported -- req.seed passed straight through (lattice_serve.rs:225) |
temperature |
supported -- validated to [0.0, 2.0] by the shared request contract, default 0.7 |
supported -- validated to [0.0, 2.0] by the same shared request contract, with the server-configured default |
top_p |
supported -- validated to (0.0, 1.0] by the shared request contract, default 0.9 |
supported -- validated to (0.0, 1.0] by the same shared request contract, with the server-configured default |
top_k |
accepted-and-ignored by ServeProfile::lattice (matches the pre-shared-contract DTO, which didn't model the field at all); generation keeps the fixed default of 50 |
supported -- mapped by ServeProfile::lattice_serve, with the server-wide default when omitted |
repetition_penalty |
accepted-and-ignored by ServeProfile::lattice (matches the pre-shared-contract DTO, which didn't model the field at all); generation keeps the fixed default of 1.1 |
supported -- mapped by ServeProfile::lattice_serve, with the server-wide default when omitted |
min_p |
engine-supported through SamplingConfig and canonical GenerateConfig; the HTTP request DTO does not expose it yet, so this surface uses the disabled default 0.0 |
same engine support and disabled server default; per-request HTTP wiring remains a later #587 slice |
logprobs / top_logprobs |
supported, non-streaming only, CPU backend -- logprobs: true + top_logprobs (0-20), rejected if combined with stream: true (shared normalize_logprobs, crates/inference/src/serve/contract.rs); shipped in PR #620, closing #585. On the Metal/Q4 backend the request passes validation but fails with a generic 500 once generation starts: the shared cross-turn-cache-aware path's check_logprobs_not_set guard rejects it (see docs/serve-http-api.md "logprobs") |
not exposed -- ChatReq models the fields (#656) only to reject them explicitly: logprobs: true/top_logprobs gets HTTP 400 "logprobs/top_logprobs are not supported by this server" (shared reject_unsupported, crates/inference/src/serve/contract.rs) rather than being silently ignored; build_cfg still hardcodes logprobs: None for the requests that do proceed (lattice_serve.rs:878) |
tools / tool_choice |
rejected with 400 if present, not silently dropped (shared reject_unsupported, crates/inference/src/serve/contract.rs) |
rejected with 400 if present (#656 closed the prior silent-ignore gap) -- ChatReq now models both fields explicitly so the shared reject_unsupported (crates/inference/src/serve/contract.rs) can name them: "tools and tool_choice are not supported by this server" |
n (multiple completions) |
rejected with 400 if > 1 (shared reject_unsupported, crates/inference/src/serve/contract.rs) |
rejected with 400 if > 1 (#656) -- ChatReq.n is modeled explicitly and checked by the shared reject_unsupported (crates/inference/src/serve/contract.rs), matching lattice serve's contract |
| Non-text content parts (image/audio/file) | one user-message image_url part is supported for non-streaming requests when the loaded Metal checkpoint passes exact vision-config/token/inventory preflight; PNG/JPEG data: URIs only, 48,000 decoded bytes maximum, remote URLs rejected. Text-only/CPU models return 400 vision_unsupported; audio/file remain unsupported. |
identical shared-contract behavior: vision-capable Metal checkpoints accept the same bounded, non-streaming inline PNG/JPEG shape, while text-only models return the same 400 code; audio/file remain unsupported. |
Unknown/malformed role |
rejected with 400, both for named-but-unsupported roles ("tool", "developer") and any other unrecognized role; production runs through shared serve::contract::normalize_request, which produces contract-owned NormalizedChatMessage values consumed by the single serve::into_engine_chat_messages adapter (see docs/serve-http-api.md "Request fields") |
rejected with the same shared contract and error codes (#641 closed by PR #656); parse_chat_req performs bounded structural JSON parsing, then normalize_request validates roles into NormalizedChatMessage before the single contract-to-engine adapter runs; the remaining local MessageRole is #[cfg(test)]-only historical error-contract coverage, not a production conversion path |
model field validation |
rejected with 400 if it doesn't match the served model id (lattice.rs:3291-) |
rejected with 400 model_not_found if it doesn't match the served model id (lattice_serve.rs, cm_lattice_serve_model_mismatch_rejected) |
| Max context / token-budget cap | max_tokens_cap hardcoded to 4096 in lattice serve's main() (not a CLI flag); Metal/Q4 backend additionally caps total context at MetalChatBackend::MAX_CACHE_LEN (4096) regardless of the model's configured max (docs/serve-http-api.md) |
supported, derived from model config (#551, closed) -- AppState.model_max_context is set at load time from the model's max_position_embeddings (model_context_from_config, lattice_serve.rs:828-), not a hard-coded constant; FALLBACK_MODEL_MAX_CONTEXT (4096, lattice_serve.rs:178) is used only when no config is available |
| Auth / rate limiting | not implemented -- no Authorization handling, no request-count or concurrency-limiting middleware (docs/serve-http-api.md) |
not implemented -- same absence, no auth/rate-limit layer in the router shown above |
| Admission control (Metal pending-job cap) | supported -- both binaries share the same MetalWorkerClient/DEFAULT_MAX_PENDING_JOBS = 32 default (crates/inference/src/serve/metal_worker.rs, issue #932), overridable per-process via --max-pending (see docs/serve-http-api.md), rejecting with HTTP 503 server_busy at the worker's admission boundary before any Metal generation work happens -- the HTTP handler has already tokenized the rendered prompt during request preparation ahead of that point, so this is not an end-to-end pre-tokenization guarantee; applies to both streaming and non-streaming requests, checked before an SSE stream is committed to on the streaming path. Not a rate limiter (see "Auth / rate limiting" above) -- bounds outstanding work on the one shared worker thread, not requests per client or unit time. No cap on the CPU backend |
supported -- same shared MetalWorkerClient admission cap (issue #932), same --max-pending override |
Capabilities present in lattice serve (the CLI's HTTP subcommand) but missing from
lattice_serve: logprobs/top_logprobs is still not implemented on lattice_serve (the field
is now explicitly rejected with 400 rather than silently ignored -- #656 -- but it doesn't run).
As of #831, stop sequences are no longer in this list -- ServeProfile::lattice_serve now
accepts and parses stop too (see the matrix row above). /v1/embeddings is no longer in this
list either: both surfaces now register the route, but the two implementations are independent --
lattice serve pools text and inline-data-URI images through a loaded vision-language checkpoint,
while lattice_serve (as of #584) embeds text only, through a separately loaded BertModel (see
the matrix row above and the HTTP routes section for both).
Capabilities present in lattice_serve but missing from lattice serve: per-request
top_k/repetition_penalty control. As of PR #619, cross-turn/prefix KV cache reuse over HTTP
is no longer in this list -- lattice serve's Metal worker now calls the cache-aware generation
path too (see the matrix row above). As of #831, static reasoning-budget forcing
(reasoning_budget request field) is no longer
in this list -- ServeProfile::lattice now accepts and applies it too (see the matrix row above).
As of ADR-080 C2
(#786), /v1/models, GET /, and disconnect-cancellation on the streaming path are no longer in
this list -- lattice serve now has all three, built on the same shared
lattice_inference::serve contract lattice_serve uses.
The shared contract also owns the Metal worker lifecycle (#833). In both binaries, each client's
Drop explicitly closes its job sender before automatic field destruction can release its owner
clone. The final owner transfers the worker's join handle to a detached reaper and waits up to two
seconds for its result. That deadline covers the full join, including thread-local destructors; if
the worker or a destructor is still running when it expires, the reaper remains detached so process
shutdown cannot hang indefinitely. Both binaries use the same server runner, which gives tracked
HTTP/1 connections up to five seconds to drain on SIGINT and Unix SIGTERM. After any drain timeout
it aborts remaining connection tasks, allows up to three seconds for cancellation cleanup, and then
exits with status 1 even if cleanup completed. That hard-exit fallback can truncate in-flight
responses, leave files partially written, and discard unflushed telemetry because it skips Rust
destructors.
As of #656, tools/tool_choice/n/non-text response_format/unknown-role/non-text-content-part
rejection now behave identically on both HTTP surfaces (all 400, none silently ignored or coerced) --
this was previously the biggest divergence between the two surfaces and it has since closed.
lattice_serve now returns HTTP 400 invalid_temperature for temperature values below 0.0
or above 2.0, and HTTP 400 invalid_top_p for top_p values at or below 0.0 or above 1.0.
These requests previously passed daemon validation and could reach generation. A top-level field
outside the shared request DTO (e.g. presence_penalty, frequency_penalty, logit_bias, user)
is ignored rather than rejected, matching standard OpenAI-compatible server behavior. The daemon's
stop, model-name, max-token clamping, top_k, repetition_penalty, reasoning-budget,
structured-output, and logprobs policies are unchanged.
Capabilities missing from both: grammar/JSON-schema constrained output over HTTP (#588),
a real /metrics endpoint (#583),
adaptive/entropy-gated reasoning budget (#500), per-request min_p wiring (engine support is the
first #587 slice). LoRA residency and per-request selection are available over HTTP on Metal.
Gaps tracked for this document: #641 (closed by PR #656) -- lattice_serve
silently coerced unknown roles to "user" and dropped non-text content parts, where lattice serve
rejected both with HTTP 400.
Every row in the Capability matrix above must map to a #[test] fn (a fixture ID, cited by
name) in the cited surface's own binary (or, for behavior shared by both binaries through the
lattice_inference::serve::metal_worker owner, that shared module -- #832), or carry an explicit
NOT-TESTABLE-YET justification. scripts/check-capability-matrix.sh (wired into
make lint-docs) greps crates/inference/src/bin/{lattice,lattice_serve,chat_metal}.rs and
crates/inference/src/serve/metal_worker.rs for fn <fixture id> and fails the build if a cited
ID no longer exists, or if a row has no fixture cell at all -- so a row that goes stale (a fixture
renamed/removed without updating this table) fails closed instead of silently drifting.
Fixtures test the request-validation/acceptance contract each row describes (a 400 with a given
code, a value resolving through to GenerateConfig, a request reaching the job queue) without a
loaded model or live GPU -- both binaries' handlers run every check in this table before they ever
touch the model (see validate_chat_request in lattice.rs, and the Job-queue split in
lattice_serve.rs). Actual token generation, real Metal dispatch, and CLI-tool output content are
out of scope for this harness and marked NOT-TESTABLE-YET; those are covered instead by the e2e
parity gate (docs/../.github/workflows/e2e-parity.yml) and manual lattice doctor/ppl_metal
runs.
| Capability row | lattice serve fixture (lattice.rs) |
lattice_serve fixture (lattice_serve.rs) |
|---|---|---|
| Chat completion (single/multi-turn) | cm_serve_model_match_passes_model_check; request-shape guards chat_completions_rejects_message_flood (MAX_MESSAGE_COUNT bound), chat_completions_matches_shared_parity_table; config-capture chat_completions_non_streaming_observation_captures_real_config_and_prompt, chat_completions_non_streaming_observation_captures_real_stopped_false |
cm_lattice_serve_model_mismatch_rejected; message-content parsing message_content_plain_string_normalizes, message_content_parts_concatenate_in_order; chat_completions_matches_shared_parity_table; config-capture chat_completions_non_streaming_observation_captures_real_config_and_messages, chat_completions_non_streaming_observation_captures_real_stopped_false; finish_reason mapping chat_completions_non_streaming_finish_reason_stop_when_engine_stopped, chat_completions_non_streaming_finish_reason_length_when_not_stopped, chat_completions_streaming_finish_reason_stop_when_engine_stopped, chat_completions_streaming_finish_reason_length_when_not_stopped |
| Streaming (SSE) | NOT-TESTABLE-YET: stream: true only changes which response arm the handler takes; the SSE body itself needs a running model to produce tokens (reject_unsupported_stream_true_ok/reject_unsupported_stream_false_ok cover the field's acceptance, not the stream). Mid-stream failure envelope IS fixtured via a fake-generate seam: chat_completions_streaming_failure_emits_error_event; shared metal_worker classification underlying it: generation_failure_is_reported_as_failed_not_complete |
NOT-TESTABLE-YET: same reason; Phase/SSE framing needs live worker output. Mid-stream failure envelope: chat_completions_streaming_failure_emits_error_event, chat_completions_streaming_failure_records_failed_metric |
| Disconnect cancellation (mid-stream) | chat_completions_streaming_disconnect_stops_generation (HTTP-level, test-utils feature: drives the real chat_completions -> body_stream -> generate_streaming_with_cancel composition against a tiny CPU model via BodyExt::frame); generate_streaming_with_cancel_true_before_prefill_returns_interrupt, generate_streaming_with_cancel_true_after_prefill_returns_interrupt, generate_streaming_with_cancel_mid_decode_stops_early_fast_path, generate_streaming_with_cancel_mid_decode_stops_early_stop_string_path (engine-level, model/qwen35/generation.rs, one fixture per checkpoint); chat_completions_streaming_disconnect_cancellation_reaches_generator_post_drop (proves the cancel signal reaches the generator closure itself, not just the HTTP layer) |
queued_job_cancelled_before_dequeue_sends_exactly_one_cancelled_event, running_job_cancelled_midstream_stops_early_and_worker_survives, running_job_cancelled_during_prefill_like_phase_never_calls_on_token (issue #832: moved to the shared lattice_inference::serve::metal_worker owner both lattice_serve.rs and, once its own migration lands on this branch, lattice.rs's Metal backend route through) |
| Cross-turn KV / GDN prefix cache | NOT-TESTABLE-YET: requires a loaded model + live Metal GPU; tracked by #462 | NOT-TESTABLE-YET: same -- the cache-aware call path (lattice_serve.rs:989-) needs a live MetalQwen35State; covered by manual/e2e verification, not this harness |
| Grammar / JSON-schema constrained output | reject_unsupported_response_format_json, reject_unsupported_response_format_text_ok |
chat_completions_json_response_format_400; structured-output validation chat_completions_rejects_duplicate_root_type_400, chat_completions_rejects_duplicate_nested_properties_400, chat_completions_rejects_duplicate_key_inside_property_schema_400, chat_completions_structured_missing_strict_400, chat_completions_structured_strict_false_400, chat_completions_structured_streaming_400, chat_completions_structured_pattern_keyword_400, chat_completions_structured_enum_keyword_400, chat_completions_structured_applies_grammar_and_marker, chat_completions_structured_blocked_constraint_500, chat_completions_structured_failed_with_blocked_wording_stays_internal_error, chat_completions_structured_length_limit_500, chat_completions_structured_validation_failed_500 |
| LoRA adapters | a_cpu_backend_refuses_an_adapter_by_name, lora_unload_on_a_cpu_backend_refuses_by_name, lora_list_on_cpu_refuses_by_name, lora_selection_on_cpu_refuses_by_name, lora_duplicate_id_is_400_before_unified_admission. Shared registry: residents_are_listed_without_applying_and_ids_are_never_reused, same_mixture_blends_once_changed_order_or_scale_reblends, absent_selection_restores_base_output_with_residents_present. |
lora_unknown_id_is_400_before_streaming, lora_duplicate_id_is_400_before_admission, lora_list_reads_confirmed_index, lora_selection_reaches_worker, lora_load_nonexistent_file_400, lora_unload_without_a_worker_is_503. |
| MTP / speculative decoding | NOT-TESTABLE-YET: enable_mtp is env-var-only, not a request field; nothing HTTP-shaped to fixture |
NOT-TESTABLE-YET: same |
| Reasoning budget / s1 (static forcing) | both_profiles_apply_reasoning_budget, context_window_accounts_for_reasoning_budget (shared serve::contract unit tests), chat_completions_non_streaming_observation_captures_real_reasoning_budget (config-capture, proves the field reaches the real GenerateConfig) |
NOT-TESTABLE-YET: reasoning_budget clamping needs build_cfg end-to-end against a real model_max_context; only build_cfg_clamps_to_runtime_context exercises the clamp math, not the full request path |
| Reasoning budget / s1 (adaptive, entropy-gated) | NOT-TESTABLE-YET: issue #500, not implemented on either surface | NOT-TESTABLE-YET: issue #500 |
| Embeddings | router-level coverage exists (embeddings_route module, lattice.rs's serve bin, #[cfg(feature = "test-utils")]): no_embedding_model_loaded_fails_closed, happy_path_text_only, happy_path_image, mixed_batch_preserves_input_order; these names fall outside this script's fixture-ID prefix convention (see the prefix list at the top of scripts/check-capability-matrix.sh), so they are not mechanically cross-checked by it, but the tests are real and pass in CI. The separate embed CLI binary/crate (text/image in, vector out) is out of this harness's scope entirely |
partially fixtured (#584): request parsing/validation/response-building is pure and metal-gpu-independent, covered by unit tests in lattice_inference::serve::embeddings (none of the fixture-ID prefixes this script tracks apply to that module's names, so it is not cited by prefix below); the HTTP handler, AppState wiring, and tests in lattice_serve.rs (embeddings_missing_input_400, embeddings_no_model_loaded_503, and others -- see the module itself) live in this binary's metal-gpu-gated imp module like the rest of its tests, so NOT-TESTABLE-YET in that harness sense, same as every other row gated the same way |
| Perplexity evaluation | NOT-TESTABLE-YET: offline CLI tool (eval_perplexity/ppl_metal), not a request-shaped capability |
not applicable |
| Quantization (Q4/QuaRot) workflow | NOT-TESTABLE-YET: offline CLI tool, not a serving-time capability | not applicable |
| Metrics / observability | NOT-TESTABLE-YET: no /metrics route on either surface (issue #583) |
NOT-TESTABLE-YET: emit_serve_event's stdout format is a Studio-app-internal wire contract, not yet fixtured here |
/v1/models |
NOT-TESTABLE-YET: list_models returns a static single-entry list from AppState.model_id via the shared models_list_body helper; route now exists (ADR-080 C2, #786) but not yet fixtured at the router level beyond the shared root_body_shape/models_list_body_shape unit tests in serve/mod.rs |
NOT-TESTABLE-YET: list_models returns a static single-entry list from AppState.model_id; trivial but not yet fixtured, tracked as a follow-up rather than blocking #654 |
GET /health |
not applicable -- trivial liveness route, no request-shaped contract to fixture | not applicable, same reason |
stop sequences |
parse_stop_strings_null_gives_empty, parse_stop_strings_single_string_gives_vec_of_one, parse_stop_strings_array_of_two_accepted, parse_stop_strings_empty_array_rejected, parse_stop_strings_array_over_four_rejected, parse_stop_strings_array_with_number_rejected, parse_stop_strings_empty_string_element_rejected, parse_stop_strings_empty_string_scalar_rejected, parse_stop_strings_array_exactly_four_accepted (9 fixtures), reject_unsupported_stop_now_accepted, cm_serve_stop_sequences_accepted_end_to_end |
chat_completions_stop_is_accepted_and_reaches_generate_config (real-worker, config-capture) |
seed |
NOT-TESTABLE-YET: pass-through with no validation logic to fixture | NOT-TESTABLE-YET: same |
temperature |
shared parity rows temperature_boundary_zero_accepted, temperature_boundary_two_accepted, temperature_out_of_range_rejected; shared contract unit tests; validate_temperature_rejects_negative, validate_temperature_rejects_above_two, validate_temperature_accepts_boundary, validate_temperature_none_uses_default |
same shared parity rows through the daemon router |
top_p |
shared parity rows top_p_boundary_one_accepted, top_p_zero_rejected, top_p_above_one_rejected; shared contract unit tests; validate_top_p_rejects_zero, validate_top_p_rejects_above_one, validate_top_p_accepts_one, validate_top_p_none_uses_default |
same shared parity rows through the daemon router |
top_k |
shared profile unit tests (profiles_preserve_sampling_extension_policy, lattice_profile_tolerates_malformed_ignored_sampling_fields -- accepted-and-ignored, not rejected) |
build_cfg_clamps_to_runtime_context plus shared profile unit tests (honoring_profile_still_rejects_malformed_sampling_fields) |
repetition_penalty |
shared profile unit tests (profiles_preserve_sampling_extension_policy, lattice_profile_tolerates_malformed_ignored_sampling_fields -- accepted-and-ignored, not rejected) |
shared profile unit tests (honoring_profile_still_rejects_malformed_sampling_fields) |
min_p |
engine tests min_p_filters_by_probability_relative_to_maximum, min_p_runs_before_top_p_and_renormalizes_survivors, and min_p_is_active_and_identical_across_cpu_sampling_paths; no HTTP request field yet |
same engine tests plus compact_metal_candidate_path_applies_min_p; no HTTP request field yet |
logprobs / top_logprobs |
validate_logprobs_absent_disables_capture, validate_logprobs_false_disables_capture, validate_logprobs_true_no_top_logprobs_defaults_to_zero, validate_logprobs_true_with_top_logprobs_ok, validate_logprobs_top_logprobs_at_boundary_twenty_ok, validate_logprobs_top_logprobs_over_twenty_rejected, validate_logprobs_top_logprobs_without_logprobs_true_rejected, validate_logprobs_top_logprobs_with_logprobs_false_rejected (8 fixtures), cm_serve_logprobs_resolved_end_to_end; unsupported-combo rejection reject_unsupported_stream_and_logprobs_rejected, reject_unsupported_logprobs_true_ok, reject_unsupported_logprobs_false_ok |
chat_completions_logprobs_400 |
tools / tool_choice |
reject_unsupported_tools_rejected, reject_unsupported_tool_choice_rejected |
chat_completions_tools_400, chat_completions_tool_choice_400 |
n (multiple completions) |
reject_unsupported_n_gt_1, reject_unsupported_n_1_ok |
chat_completions_n_greater_than_one_400 |
| Non-text content parts (image/audio/file) | vision_content_parts_1135::{vision_model_accepts_image_and_enqueues_it_on_the_shared_worker,text_only_model_rejects_the_same_image_with_capability_code}, shared contract/adapter/worker tests |
chat_completions_vision_model_enqueues_image_on_shared_worker, chat_completions_image_url_400, chat_completions_oversized_part_400, shared contract/adapter/worker tests; model-gated production route: vision_serve_e2e_test; content-part rejection message_content_image_url_rejected, message_content_unknown_part_rejected |
Unknown/malformed role |
to_chat_messages_rejects_invalid_role, to_chat_messages_rejects_tool_role, to_chat_messages_rejects_developer_role |
chat_completions_unknown_role_400, message_role_unknown_rejected, message_role_tool_and_developer_rejected_as_unsupported_feature, chat_completions_tool_and_developer_role_400_unsupported_feature |
model field validation |
cm_serve_model_mismatch_rejected, cm_serve_model_match_passes_model_check |
cm_lattice_serve_model_mismatch_rejected |
| Max context / token-budget cap | validate_max_tokens_rejects_zero, validate_max_tokens_rejects_above_cap, validate_max_tokens_alias_conflict_rejected, validate_max_tokens_at_exactly_cap_ok, validate_max_tokens_uses_default_when_absent, validate_max_tokens_alias_agrees, validate_max_tokens_max_completion_only_ok; context_window_accepts_boundary_and_rejects_overflow, context_window_accounts_for_reasoning_budget (#831: shared saturating full-window formula, prompt + max_new_tokens + reasoning_budget + 1 <= max_context, used by both binaries' preflight); chat_completions_streaming_context_overflow_returns_400_before_committing_sse |
build_cfg_clamps_to_runtime_context, chat_completions_conflicting_max_tokens_400, model_context_uses_config_max_position_embeddings, model_context_falls_back_to_4096_when_absent, check_prompt_fits_window_rejects_when_prompt_plus_decode_overflows, check_prompt_fits_window_accepts_ordinary_prompt_unclamped (2 fixtures), chat_completions_max_tokens_zero_400, chat_completions_streaming_context_overflow_returns_400_before_committing_sse, chat_completions_streaming_context_overflow_matches_lattice_real_router, build_cfg_aliases_max_completion_tokens_when_max_tokens_absent |
| Auth / rate limiting | not applicable -- no auth/rate-limit layer exists on either surface to fixture | not applicable, same |
| Admission control (Metal pending-job cap) | chat_completions_streaming_returns_503_before_sse_commit_at_cap, chat_completions_non_streaming_returns_503_server_busy_not_500_at_cap |
chat_completions_returns_503_json_envelope_when_admission_cap_reached |
| Request precedence / order | cm_serve_unsupported_feature_rejected_before_model_check, cm_serve_empty_messages_rejected, cm_serve_last_message_not_user_rejected (previously untested: these three checks ran only inline in chat_completions with no dedicated fixture at all); cm_serve_context_window_checked_before_stop_parsing pins that the context-window preflight runs before stop-sequence parsing, restoring the original inline order across the prepare_chat_request extraction |
shared normalize_request runs unsupported-feature checks before model/message/sampling validation; parse_chat_req only performs bounded structural JSON parsing -- bounds enforced by parse_chat_req_rejects_too_many_parts_before_typed_parse, parse_chat_req_accepts_max_parts_boundary, parse_chat_req_rejects_oversized_text_part_before_typed_parse. The chat_completions_*_400 fixtures above cover this ordering, and the shared context-precedence fixture pins context validation before stop parsing |
chat_metal is not part of the two-column OpenAI-shaped matrix above -- its --json --serve
mode speaks a simpler newline-delimited-JSON stdin protocol ({"prompt":...} in,
@@lattice {"ev":"gen_token",...} out), not the /v1/chat/completions wire shape, and it has
no stop/messages/tools/role fields at all: the client resends the full ChatML-rendered
history as a single prompt string every request (see the module doc comment at the top of
chat_metal.rs). Its request-parsing was previously inline in the serve loop with no tests at
all; #654 factored it into parse_serve_request_line (chat_metal.rs) specifically so it could
be fixtured:
| Wire behavior | Fixture |
|---|---|
prompt-only request falls back to all CLI defaults |
cm_chat_metal_serve_prompt_only_uses_all_defaults |
Per-request field overrides (max_tokens/temperature/top_k/top_p/repetition_penalty/seed/reasoning_budget) |
cm_chat_metal_serve_per_request_overrides_all_fields |
| Malformed JSON line | cm_chat_metal_serve_malformed_json_rejected |
Missing prompt field |
cm_chat_metal_serve_missing_prompt_rejected |
reasoning_budget: 0 treated as absent (not an explicit zero-budget override) |
cm_chat_metal_serve_zero_reasoning_budget_falls_back_to_default |
No stop/history fields exist on the wire shape -- an extra stop key is inert, not an error |
cm_chat_metal_serve_stateless_no_stop_strings_or_history_fields |
Actual Metal generation, @@lattice gen_token event stream content, cross-turn cache reuse behavior |
NOT-TESTABLE-YET: requires a loaded model + live Metal GPU |
docs/serve-http-api.md-- fulllattice serveHTTP API reference (request/response shapes, validation order, error envelopes, streaming details)docs/cross-turn-cache.md-- cross-turn KV/GDN prefix cache: what calls it and what doesn'tdocs/q4-quantization.md-- Q4/QuaRot quantization workflow andlattice doctoroutput- #604 -- usage-example/cookbook coverage tracking, including #601 (
lattice_serveHTTP API walkthrough) and #602 (CLI tool walkthroughs)