From 9d6b3da1e4bba3bf0613543f2575979bc14d7aaf Mon Sep 17 00:00:00 2001 From: Chris Phillipson Date: Mon, 27 Jul 2026 09:28:19 -0700 Subject: [PATCH 1/8] =?UTF-8?q?docs:=20ADR-0011=20=E2=80=94=20local=20mode?= =?UTF-8?q?ls,=20out-of-band=20provenance,=20$0=20per=20model?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The usage scorecard assumes every transcript came from a metered vendor endpoint. `ollama launch claude` / `ollama launch codex` makes that assumption a one-command violation, and nothing in the usage path knows ollama exists (it appears only in routing.mjs/providers.mjs today). Research findings, cited in the ADR: - The model id is echoed verbatim by Ollama's Anthropic compat layer, so identity capture already works — but the tag carries no quantization in the general case (qwen3.6:latest, qwen3:32b), which lives only in /api/tags details{} and /api/show model_info{}. - /v1/messages returns usage.{input_tokens,output_tokens} only; prompt caching is explicitly unsupported. Every local session therefore records cacheRead=0, pricing 100% of input at 1x — the inverse of the 96.3% cache-read corpus ADR-0009 §3 is built on. - Ollama documents `ollama cp qwen3-coder claude-3-5-sonnet` for tools that hardcode vendor names, so the model id is forgeable by a recommended workflow. Id-based local detection cannot work. - Measured: 754 indexed files, 7 distinct model ids, all matching the price table — FALLBACK_PRICE has never fired. The silent path costs nothing today and would misprice every local turn tomorrow. Decisions: provenance established out-of-band (loopback catalogue read plus ak's own wiring), ambiguous when a local tag collides with a table entry; three cost states (metered / local $0 / unpriced) retiring FALLBACK_PRICE from the cost path; per-model rows for local models rather than one bucket; digest-keyed identity displayed with quantization; token KPI split the same way cost is; provider capability gaps stated per session. Status is Proposed, not Accepted: no local session exists on this machine to observe, so the transcript-side claims are derived from specs rather than measured. docs/LOCAL-MODEL-VALIDATION.md is the protocol that closes that gap — six questions with predicted answers, a content-free capture step, and a results table naming the ADR action for each contradiction. Also amends ADR-0009 §3 and §8, and adds the missing ADR-0010 row to the ADR index alongside 0011. --- docs/LOCAL-MODEL-VALIDATION.md | 191 +++++++++++ ...ge-scorecard-local-transcript-analytics.md | 18 ++ ...nance-zero-cost-and-transcript-fidelity.md | 300 ++++++++++++++++++ docs/adr/README.md | 12 +- 4 files changed, 520 insertions(+), 1 deletion(-) create mode 100644 docs/LOCAL-MODEL-VALIDATION.md create mode 100644 docs/adr/0011-local-model-provenance-zero-cost-and-transcript-fidelity.md diff --git a/docs/LOCAL-MODEL-VALIDATION.md b/docs/LOCAL-MODEL-VALIDATION.md new file mode 100644 index 00000000..73c60a81 --- /dev/null +++ b/docs/LOCAL-MODEL-VALIDATION.md @@ -0,0 +1,191 @@ +# Validating local-model capture: a two-session protocol + +**For:** whoever runs the evidence pass that moves +[ADR-0011](adr/0011-local-model-provenance-zero-cost-and-transcript-fidelity.md) from **Proposed** +to **Accepted**. + +**Why this exists.** ADR-0011 decides how the usage scorecard should treat sessions served by a +**local** model (`ollama launch claude`, `ollama launch codex`). Every claim it makes about *what +Claude Code and Codex actually write to disk* in that situation is **derived from vendor +specifications, not observed** — at research time (2026-07-27) the Ollama daemon was down and all 754 +indexed transcripts on the reference machine were vendor-metered. This document is the experiment +that replaces derivation with measurement. + +**Time:** about 15 minutes of interactive work, plus two paste-backs. + +--- + +## What we are trying to find out + +Six questions. The first four decide whether ADR-0011's sections are right; the last two decide +whether two of its stated limitations are real. + +| # | Question | ADR-0011 depends on it in | Predicted answer | +|---|---|---|---| +| Q1 | What literal string lands in `message.model` for a local Claude Code session? | F1, §1 | the Ollama tag, verbatim (e.g. `qwen3.6:latest`) | +| Q2 | Does `message.usage` carry `cache_read_input_tokens` / `cache_creation_input_tokens`? | F3, §3 | **no** — absent, so the indexer coerces both to `0` | +| Q3 | Is an `ai-title` event emitted when titling is served locally? | F7a, §7 | uncertain — this is the least-predictable answer | +| Q4 | Does a **mid-stream interruption** produce an `isApiErrorMessage` turn? | F7b, §7 | uncertain — if not, local failures under-report | +| Q5 | Does a *second* turn's `input_tokens` behave like a full prompt count, or a KV-cache delta? | F4, §6 | unstable; may drop sharply on turn 2 | +| Q6 | Does a model id aliased onto a vendor name reach the transcript intact? | **F5, §1 — the load-bearing one** | yes — `claude-opus-5` written verbatim | + +Q6 is optional but worth more than the rest combined: ADR-0011's entire out-of-band provenance design +exists **because** the model id is forgeable. Right now that rests on a documentation sentence. + +--- + +## Before you start + +Leave the Ollama daemon **running** for the whole exercise, including afterwards — the capture step +reads its catalogue. + +```bash +ollama serve & # or just launch anything; skip if already running +ollama --version +curl -s http://localhost:11434/api/tags | head -c 200 # must return JSON, not exit 7 +``` + +Record the versions you are testing (they belong in the results): + +```bash +ollama --version && claude --version && codex --version +``` + +--- + +## Session 1 — Claude Code (answers Q1–Q5) + +Use a tag whose **quantization is not in the name**, so the run also exercises the F2 metadata gap: + +```bash +ollama launch claude --model qwen3.6:latest +``` + +Then, inside the session, in this order: + +1. **Force a tool call.** *"Read package.json and tell me the version field."* + → confirms `tool_use` blocks survive, which the tool-mix classifier prior depends on. +2. **Two or three more ordinary prompts.** *"What test runner does this project use?"*, *"Summarize + the scripts section."* One turn cannot answer Q5 — the second turn is the one that shows whether + a warm prefix changes the reported input count. +3. **Let it sit long enough for a session title to appear** (Q3). +4. **On the final response, press `Ctrl-C` mid-stream** — while tokens are still printing, not after + it finishes. This is Q4, and it is the only question no amount of reading can answer. + +Note the **working directory** you ran in; the capture step needs it. + +## Session 2 — Codex (confirms the second host) + +```bash +ollama launch codex +``` + +Two or three prompts, at least one forcing a file read. This confirms `turn_context.model` carries +the local tag, that `token_count` events still appear (Codex meters locally, so they should), and +that `rate_limits` is **absent** rather than zeroed — a zeroed rate limit would render as "0% of plan +used", which would be a fabricated denominator of exactly the kind ADR-0010 forbids. + +## Optional — the alias experiment (Q6) + +```bash +ollama cp qwen3-coder:30b claude-opus-5 +ollama launch claude --model claude-opus-5 # one prompt is plenty +ollama rm claude-opus-5 +``` + +Two things to know before running it: + +- It leaves **one session in your corpus that today's code prices as genuine Opus at $5/$25 per 1M**. + That is the demonstration, not a side effect. Note the session so it can be identified later. +- `ollama rm` removes only the alias, not the underlying weights. Your model store is untouched. + +--- + +## Capture (run after both sessions) + +Substitute the working directory you used. The commands below print **field names, counts, and model +ids only — never message content** — so the output is safe to paste into an issue or a PR. + +```bash +# 1. Locate the two most recent Claude transcripts +ls -t ~/.claude/projects/*/*.jsonl | head -3 + +# 2. Q1 — every model id the session recorded +T=$(ls -t ~/.claude/projects/*/*.jsonl | head -1) +grep -o '"model":"[^"]*"' "$T" | sort | uniq -c + +# 3. Q2/Q5 — the usage object, verbatim, per assistant turn +grep -o '"usage":{[^}]*}' "$T" + +# 4. Q3 — was a title event written? +grep -c '"type":"ai-title"' "$T" + +# 5. Q4 — did the interrupted stream produce an error turn? +grep -c '"isApiErrorMessage":true' "$T" + +# 6. Codex side +C=$(ls -t ~/.codex/sessions/**/rollout-*.jsonl 2>/dev/null | head -1) +grep -o '"model":"[^"]*"' "$C" | sort | uniq -c +grep -c '"type":"token_count"' "$C" +grep -c 'rate_limits' "$C" + +# 7. The catalogue, for the §5 identity claims (digest + quantization) +curl -s http://localhost:11434/api/tags \ + | python3 -c 'import json,sys; [print(m["name"], m["digest"][:12], m["details"].get("parameter_size"), m["details"].get("quantization_level"), m["details"].get("format")) for m in json.load(sys.stdin)["models"]]' + +# 8. Full metadata for the one tag you tested +curl -s http://localhost:11434/api/show -d '{"model":"qwen3.6:latest"}' \ + | python3 -c 'import json,sys; d=json.load(sys.stdin); print(d.get("capabilities")); print({k:v for k,v in d.get("model_info",{}).items() if "context_length" in k or k.startswith("general.")})' +``` + +--- + +## Results + +Fill this in and it becomes the record. An answer that **contradicts** the prediction is the valuable +outcome — it means ADR-0011 gets corrected before any code is written against it. + +| # | Predicted | Observed | ADR action if it differs | +|---|---|---|---| +| Q1 | tag verbatim | | F1 and §1's catalogue-membership check need rework | +| Q2 | no cache fields | | if present, §3's premise weakens and §7's fidelity note narrows | +| Q3 | uncertain | | if titles are emitted, §7 drops the titling caveat | +| Q4 | uncertain | | if no error turn, §7 must state that local failures under-report | +| Q5 | unstable | | if stable, §6's token-KPI split may be unnecessary | +| Q6 | id verbatim | | **if the id is rewritten, §1 can be simplified drastically** | + +| Environment | Value | +|---|---| +| `ollama --version` | | +| `claude --version` | | +| `codex --version` | | +| date run | | +| tag tested | | + +**Where the results go:** + +1. Observations and any corrected findings → [`USAGE-SCORECARD-METRICS.md`](USAGE-SCORECARD-METRICS.md), + which ADR-0011 names as the home for verifiable per-figure detail. +2. Corrections to the findings themselves → edit ADR-0011's F1–F7 in place, since it is still + **Proposed** and has not yet been built against. +3. Status flip → ADR-0011 `Proposed` → `Accepted`, and the row in + [`adr/README.md`](adr/README.md) updated to match. +4. This document → `docs/archive/2026-07-27-local-model-validation-protocol.md` with an index row + stating what it proved, per [`archive/README.md`](archive/README.md)'s naming convention. It is a + worklist; when the work is closed it is history, not a living doc. + +--- + +## Two decisions that are not the experiment's to make + +Both are recorded here so they are not silently assumed while the evidence is being gathered: + +1. **Whether `ak` may read `127.0.0.1:11434` at all.** ADR-0009 §2 promised "zero network calls"; + ADR-0011 §2 argues a loopback read is not egress and sits inside ADR-0010's provider-mediated + precedent. That reasoning is sound but it modifies a promise made in writing, so it wants a + deliberate yes rather than an inference. +2. **Whether to call `/api/show` per tag.** `/api/tags` alone yields family, parameter size, + quantization, and digest — enough for the display string in §5. `/api/show` adds architecture and + context length at one extra call per distinct tag. Recommendation: ship `/api/tags` only, and add + `/api/show` if context length proves to matter. Step 8 of the capture collects it either way, so + the decision can be made on real output. diff --git a/docs/adr/0009-usage-scorecard-local-transcript-analytics.md b/docs/adr/0009-usage-scorecard-local-transcript-analytics.md index 3dfb0fc7..1145cca6 100644 --- a/docs/adr/0009-usage-scorecard-local-transcript-analytics.md +++ b/docs/adr/0009-usage-scorecard-local-transcript-analytics.md @@ -102,6 +102,15 @@ would make such a percentage honest, and an invented denominator is worse than n OpenAI publishes no pricing in `~/.codex/models_cache.json` (verified), so Codex rates are a maintained table and are **date-stamped** in the UI so staleness is visible rather than silent. +> **Amended by [ADR-0011](0011-local-model-provenance-zero-cost-and-transcript-fidelity.md) +> (2026-07-27):** this section assumes every transcript came from a metered vendor endpoint. A +> session served by a **local** model (`ollama launch claude` / `ollama launch codex`) breaks three of +> its premises at once — no cache accounting exists to apply the 0.1× multiplier to, token counts are +> the provider's own approximations, and the model id is one the vendor documents how to alias onto a +> Claude name. ADR-0011 therefore splits cost into `metered` / `local` ($0, exact) / `unpriced`, and +> **retires `FALLBACK_PRICE` from the cost path**: pricing an unrecognised model at Sonnet-class rates +> is the invented-denominator error this section rejects, applied to a rate instead of a limit. + ### 4. Engaged time is the union of *active* intervals — three tiers, and the honest one leads Reporting summed spans would have claimed 21 h/day. Merging overlapping session spans fixes only @@ -290,6 +299,15 @@ prompt" and "not the human" are different claims. Codex rollouts record only rea `user_message` events, so every Codex user turn is `kind: prompt` by construction. Full mechanics: [`docs/TRANSCRIPTS.md`](../TRANSCRIPTS.md). +> **Amended by [ADR-0011](0011-local-model-provenance-zero-cost-and-transcript-fidelity.md) +> (2026-07-27):** this section's principle — *withheld content announces itself* — was scoped to +> masking and truncation, both things **this panel** does. A **provider** withholds too: a local +> Ollama-backed session reports no cache accounting and approximate token counts, and its titling and +> mid-stream errors are served locally, so `cacheRead: 0` on a local session is a fact about the +> provider and not about the work. ADR-0011 §7 extends the same announcement rule to provider +> capability, so a reader can tell "this workload had no cache hits" from "this provider cannot +> report cache hits" — which the panel currently renders identically. + ## Consequences ### Good diff --git a/docs/adr/0011-local-model-provenance-zero-cost-and-transcript-fidelity.md b/docs/adr/0011-local-model-provenance-zero-cost-and-transcript-fidelity.md new file mode 100644 index 00000000..8f1dfa28 --- /dev/null +++ b/docs/adr/0011-local-model-provenance-zero-cost-and-transcript-fidelity.md @@ -0,0 +1,300 @@ +# ADR-0011 — Local models: provenance out-of-band, $0 per model, and stated transcript fidelity + +- **Status:** Proposed — see *Validation required* before this may be marked Accepted +- **Date:** 2026-07-27 +- **Deciders:** agentic-kit maintainers +- **Amends:** [ADR-0009](0009-usage-scorecard-local-transcript-analytics.md) §3 (cost) and §8 (transcripts) + +## Context + +`ak` already treats `ollama` as a first-class provider on the *routing* axis: it is in +`PROVIDERS` (`src/lib/routing.mjs:30`), in `SUBSCRIPTION_PROVIDERS` as a $0 local backend +(`:36`), in the aqe provider matrix (`src/lib/providers.mjs:56`), and in the billing hint the +provider picker prints — *"ollama/onnx = local ($0)"* (`src/commands/x/provider.mjs:54`). + +The **usage** axis knows nothing about it. Grepping `ollama|OLLAMA|ANTHROPIC_BASE_URL|OPENAI_BASE_URL` +across `src/` returns hits in exactly those three routing/provider files and none in +`usage-index.mjs`, `pricing.mjs`, or `dashboard-server.mjs`. The scorecard is built on an assumption +it never states: that every transcript on disk was produced by a metered vendor endpoint. + +Ollama has since made that assumption easy to violate from the host side. `ollama launch` configures +and starts a frontier CLI against a local model — `ollama launch claude`, `ollama launch codex`, and +fourteen other integrations — so a Claude Code or Codex session backed by a local model is now a +one-command act, not a hand-wired experiment. + +**Machine evidence gathered for this ADR** (2026-07-27): `ollama` client **0.32.1** installed, daemon +**not running** at scan time (`curl :11434` → exit 7), **132 GB** of models on disk, 14 tags across +8 families — including `qwen3.6:35b-a3b-q4_K_M`, `qwen3.6:27b-mlx`, `qwen3.6:latest`, +`qwen3-coder:30b`, `qwen3-coder:30b-a3b-q4_K_M`, `gemma4:26b-a4b-it-q4_K_M`, `glm-4.7-flash:q4_K_M`, +`qwen3:32b`, and `hf.co/yuxinlu1/gemma-4-12B-coder-fable5-composer2.5-v1-GGUF:Q8_0`. + +### Findings + +Sources are cited inline; every claim below is either a vendor document, upstream source, or a +measurement on this machine. Where a claim is *derived* rather than observed, it says so. + +**F1 — Model identity already flows through, verbatim.** Ollama's Anthropic compatibility layer +(`anthropic/anthropic.go`, `ToMessagesResponse()`) **echoes the request's `model` string** into the +response ([DeepWiki: Anthropic Compatibility Layer](https://deepwiki.com/ollama/ollama/3.5-anthropic-compatibility-layer)). +Claude Code writes that value to `message.model`, which `usage-index.mjs:480` reads verbatim. So a +local session's model *name* is already captured correctly today — `qwen3.6:35b-a3b-q4_K_M` will +appear as a model in play. Nothing needs building for identity. What is missing is everything +attached to it. + +**F2 — The tag is not the model.** Quantization, parameter count, family, format, and context length +are **not in the tag** in the general case: `qwen3.6:latest`, `qwen3:32b`, and `qwen3-coder:30b` carry +no quant, while `gemma4:26b-a4b-it-q4_K_M` does. That metadata lives only in the daemon's own +catalogue — `/api/tags` returns `details.{format, family, families, parameter_size, +quantization_level}` plus `digest` and `size` ([Ollama — List models](https://docs.ollama.com/api/tags)); +`/api/show` adds `model_info` (`general.architecture`, `general.parameter_count`, `general.file_type`, +`*.context_length`) and `capabilities`; `/api/ps` adds `size_vram`, `context_length`, and `expires_at` +for what is actually resident ([Ollama — List running models](https://docs.ollama.com/api/ps); +[`docs/api.md`](https://github.com/ollama/ollama/blob/main/docs/api.md)). **Surfacing the exact model — +"qwen3.6 · 35B · Q4_K_M · gguf" — therefore requires an out-of-band read. The transcript alone cannot +answer it.** + +**F3 — There is no cache accounting, structurally.** Ollama's `/v1/messages` returns +`usage.{input_tokens, output_tokens}` and nothing else; `cache_creation_input_tokens` and +`cache_read_input_tokens` are absent, and **prompt caching is on the explicit unsupported list** +alongside the token-counting endpoint, Batches, citations, and PDFs +([Anthropic compatibility](https://docs.ollama.com/api/anthropic-compatibility)). Since +`usage-index.mjs:483-489` coerces missing counters to `0`, every local session records +`cacheRead = cacheWrite = 0` — so **100% of its input prices at 1×**. ADR-0009 §3 built the entire +cost model around the opposite fact: 96.3% of the reference corpus is cache reads billing at 0.1×. +A local session is the pathological input to that formula. + +**F4 — Local token counts are approximations, and not stable ones.** The compatibility doc states +plainly that *"token counts are approximations based on the underlying model's tokenizer."* Upstream, +`prompt_eval_count` has a long history of reporting something other than "tokens in this prompt" — +it disappears entirely on repeated identical requests ([ollama#2068](https://github.com/ollama/ollama/issues/2068)) +and has been reported outright broken ([ollama#3427](https://github.com/ollama/ollama/issues/3427)), +with KV-cache reuse meaning the count reflects what was *evaluated*, not what was *sent*. So local +figures are not comparable to vendor-metered ones **in tokens**, before any question of dollars. + +**F5 — The model id can be forged, and the vendor recommends it.** For tools that hardcode Anthropic +names, Ollama's own documentation recommends aliasing: +`ollama cp qwen3-coder claude-3-5-sonnet` (Anthropic compatibility; the same advice appears for +OpenAI names). A user following that advice with `ollama cp qwen3-coder claude-opus-5` produces +transcripts whose `message.model` matches `PRICES['claude-opus-5']` **exactly** and bills a free local +session at $5/$25 per 1M. **Any scheme that infers "is this local?" from the model string is +defeated by a documented, recommended workflow.** Provenance must come from a channel the model id +cannot forge. + +**F6 — The fallback path is currently unexercised, which is precisely why it is dangerous.** Measured +over the live index (`~/.config/agentic-kit/usage-index.json`, schema 6, **754 indexed files**), the +complete set of model ids carrying usage rows is: `claude-sonnet-5`, `claude-opus-4-8`, +`claude-fable-5`, `claude-opus-5`, `claude-opus-4-7`, `claude-haiku-4-5-20251001`, `gpt-5.6-sol` — +**seven ids, all of which match the price table**. (The bare `sonnet`/`haiku`/`opus` strings visible +in raw transcripts are `Agent`-tool *inputs*, not `message.model`, and never reach pricing — checked.) +So `FALLBACK_PRICE` ($3/$15, Sonnet-class, `matched:false`) has never once been hit on this corpus, +and `matched` has no consumer anywhere in `dashboard-server.mjs` or `usage-index.mjs`. The first +local session would be the first thing to exercise a silent path — and it would exercise it for +every turn. + +**F7 — Transcript fidelity is high on content, gapped on metadata.** `/v1/messages` supports +messages, multi-turn, streaming, system prompts, **tool calling**, **thinking**, and vision. Tool +calling is the load-bearing one: `tool_use` blocks are what `usage-index.mjs:492-497` counts, so the +tool-mix prior in `usage-classify.mjs` and the per-turn tool chips keep working. Three gaps: +*(a)* `ai-title` is generated by the host's own titling call, which under a local base URL is served +by the local model — when it yields nothing, `rec.title` falls back to `clip(firstPrompt)` +(`:507`), and classification loses its 93%-coverage signal; *(b)* **server-sent errors during +streaming are unsupported**, so a mid-stream local failure may not arrive as the `isApiErrorMessage` +turn that ADR-0009's exception accounting depends on — local failures can under-report; *(c)* on the +Codex side `ollama launch codex` writes a dedicated profile, `~/.codex/ollama-launch.config.toml`, +with `base_url = "http://localhost:11434/v1/"` and `wire_api = "responses"` +([Ollama — Codex CLI](https://docs.ollama.com/integrations/codex)) — token accounting there is +Codex's own, so `token_count` events survive, while `rate_limits` is simply absent and +`rec.rateLimits` correctly stays `null`. + +## Decision + +### 1. Provenance is established out-of-band and stamped on the session — never inferred from the model id + +F5 forecloses id-based inference. Two channels, in order: + +1. **The local catalogue.** `GET http://127.0.0.1:11434/api/tags` yields the exact set of tags and + digests this machine can serve. Membership is the primary signal. +2. **ak's own wiring.** `ANTHROPIC_BASE_URL` in the managed `.claude/settings.local.json` `env` + block, and the presence of `~/.codex/ollama-launch.config.toml` — both artifacts `ak` or + `ollama launch` wrote, both outside the transcript. + +When a model id is **both** a local tag and a price-table entry (the F5 alias case), the session is +**`ambiguous`**: no dollar figure, no $0 claim, and the ambiguity is displayed with its cause +("`claude-opus-5` is also a local Ollama tag on this machine"). Resolving it would require guessing, +and this is the ADR series' standing answer to a guess — ADR-0009 §5's `Unclassified` and §6's +*"no $ claimed"* are the same move. + +**Stated limitation, load-bearing:** catalogue membership is read **now** and the session ran +**then**. A tag pulled today does not prove last week's session was local, and a tag deleted since +does not prove it was not. Provenance is therefore stamped **at index time**, carries the evidence +that produced it (`catalog` / `config` / `both`), and is re-derivable on reindex. This is weaker +than the vendor-mediated facts of ADR-0010 and is labelled as inference in the UI, not as a fact. + +### 2. A loopback read of the local daemon is not egress — ADR-0010's pattern, applied to a third provider + +ADR-0009 §2 promised "zero network calls" and ADR-0007 drew the `dashboard`/`admin` line at +**network egress**. A request to `127.0.0.1:11434` crosses no network boundary and touches no +credential — it is the same shape as ADR-0010 spawning `codex app-server` to read quota, and the +same trust model as the dashboard shelling out to `ak status`. It stays in `dashboard`. + +Rules: read-only (`/api/tags`, and `/api/show` only for tags actually seen in transcripts); cached to +`~/.config/agentic-kit/ollama-catalog.json` with a TTL; **never starts the daemon** (it was down on +this machine during research — that must be an ordinary state, not a degraded one); on failure the +feature returns `provenance: unknown` and the panel says the catalogue was unreachable. No auto-pull, +no model loading, no write path to Ollama, ever. + +### 3. Three cost states replace one silent fallback + +`matched` graduates from an unread flag to the thing that decides what may be claimed: + +| State | Condition | Cost shown | +|---|---|---| +| `metered` | id matches the price table **and** provenance is not local | dollars, as today | +| `local` | provenance is local | **exactly `$0`** — a fact, not an estimate | +| `unpriced` | no table match, no local provenance | **no dollar figure at all** | + +`FALLBACK_PRICE` **retires from the cost path.** Pricing an unrecognised model at Sonnet-class rates +was defensible when it could only mean "the table is one release behind"; with local hosts in play it +means "we invented a rate for tokens that may have cost nothing," which is the fabricated-denominator +error ADR-0009 §3 rejects, wearing a different hat. `priceFor` keeps returning the fallback for +callers that want a *rate*; `costOf` returns `null` for `unpriced`, and `null` renders as "unpriced", +never as `$0`. **`$0` and "no figure" must never render the same** — one is a measurement, the other +is a refusal. + +### 4. Per-model rows for local models — no blanket "local" bucket + +Collapsing local usage to a single row would answer "did I use local models" and destroy "*which* +local models, and how much". Every local model keeps **its own row**, ranked by **tokens** (ranking +by cost is meaningless when every row is `$0`), carrying tokens, responses, and sessions. Above them +sits one aggregate line — *"N local models · M sessions · T tokens · $0"* — so the count is directly +readable rather than something the reader has to tally. The Scorecard's headline cost KPI gains a +sibling: metered spend and local-and-free stand side by side, never summed. + +### 5. Local model identity is keyed on digest and displayed with its quantization + +Tags are mutable (`:latest` moves; `ollama cp` makes two names for one blob), so the **digest** is +the only stable identity. Records carry +`{ tag, digest, family, parameterSize, quantization, format, contextLength }` from `/api/tags` + +`/api/show`, displayed as **`qwen3.6:35b-a3b-q4_K_M · 35B · Q4_K_M · gguf · 262k ctx`** and keyed on +digest. This is what makes the panel able to say two rows are the same weights under different names, +or that "qwen3.6" in June and "qwen3.6" in July were different models — neither of which the tag can +express. Metadata absent (daemon down, tag since removed) renders as the tag alone; a missing field +is shown missing, not defaulted. + +### 6. Token totals split the same way cost does + +F4 means local token counts are approximations from a different tokenizer with unstable +prompt-accounting semantics. Blending them into one headline "tokens" KPI would produce a number +that is neither metered-accurate nor locally-accurate. The KPI therefore reports metered and local +separately, on the same footing as §3's cost split. One honest number in two parts beats one +dishonest number. + +### 7. Transcript fidelity is stated per session, not silently degraded + +ADR-0009 §8's principle — *withheld content announces itself* — extends from truncation and masking +to **provider capability**. Local sessions render turn-for-turn identically (F7: tools, thinking, +vision, multi-turn all survive), and carry a **fidelity note** listing what the provider could not +report: + +- **no cache accounting** — `0` cache reads is a fact about Ollama, not about the work; +- **approximate token counts** — per the vendor's own wording; +- **no vendor rate limits** — the Limits view shows nothing for this session rather than a stale + Anthropic figure (ADR-0010's staleness reporting already covers the mechanism); +- **titling served locally** — when `ai-title` is absent the title is the first prompt, and the + category basis says so. + +A reader who sees `cacheRead: 0` next to a metered session showing 96% cache reads must be able to +tell "this workload had no cache hits" from "this provider cannot report cache hits". Those are +different statements and the panel currently renders them identically. + +### 8. Ollama is the implementation; the seam is "OpenAI/Anthropic-compatible local endpoint" + +LM Studio, llama.cpp's server, vLLM, and any OpenAI-compatible gateway create the same three +problems (no cache fields, forgeable ids, absent quota). Only Ollama is implemented, because it is +what `ak` already wires and what is installed here. The provenance channel in §1 is defined as an +interface — *"a catalogue that can be asked which models are local"* — so a second backend is a new +catalogue reader, not a second cost model. **Nothing else is implemented on speculation.** + +## Validation required + +This ADR is **Proposed**, not Accepted, and the reason is recorded rather than papered over: **no +local-model session exists on this machine to observe.** The Ollama daemon was down during research, +and all 754 indexed files are vendor-metered. Every claim about *what Claude Code actually writes* +for an Ollama-backed session — §3's `cacheRead: 0`, F7's title and streaming-error behaviour — is +**derived from two specifications, not measured**. + +Before this may be marked Accepted: + +1. Run one `ollama launch claude` session and one `ollama launch codex` session against a known tag. +2. Capture the resulting transcript and confirm: the literal `message.model` string; presence/absence + of `usage.cache_*`; whether an `ai-title` event is emitted; and what a deliberately interrupted + stream records. +3. Confirm `/api/tags` `digest` values match what the transcript's tag resolves to. +4. Record the observations in `docs/USAGE-SCORECARD-METRICS.md` and correct any finding above that + the evidence contradicts. + +The executable form of this list — exact commands, the six questions with their predicted answers, +the content-free capture step, and a results table — is +**[`docs/LOCAL-MODEL-VALIDATION.md`](../LOCAL-MODEL-VALIDATION.md)**. It also carries the optional +alias experiment (`ollama cp qwen3-coder claude-opus-5`) that would turn F5 from a documented claim +into a measured one, which is the single highest-value observation available: **§1 exists entirely +because the model id is forgeable, and if it turns out not to be, §1 collapses to something much +simpler.** + +## Consequences + +### Good + +- Local sessions become **visible and correctly free** rather than invisible or fictitiously + expensive, and "how many local models ran at $0" is a directly readable number (§4). +- The exact model — family, parameter size, **quantization**, context — is surfaced, which the + transcript alone can never provide (F2). +- Retiring the silent fallback (§3) fixes a latent honesty defect that predates local models: + an unrecognised **vendor** id also stops being quietly priced as Sonnet. +- The offline-first contract holds. Loopback is not egress, and the daemon being absent is an + ordinary state. + +### Costs and risks + +- **Provenance is as-of-read-time, not as-of-session-time** (§1). Explicitly labelled inference. +- **The alias case is reported, not resolved** (§1). A user who aliases a local model to a vendor + name gets `ambiguous` rows until they rename. Accepted: the alternative is guessing about money. +- **Schema bump.** Provenance and local-metadata fields force `SCHEMA_VERSION` 6 → 7 and a full + reindex (~1 min cold, per ADR-0009 §2). +- **A new dependency surface.** Ollama's API is not versioned in lockstep with `ak`; `/api/tags` + field drift degrades metadata to the tag alone rather than breaking the panel. +- **Local token counts stay approximate** (F4). The panel can label that, not fix it. +- **`$0` is right for marginal cost and silent about electricity, hardware, and time.** The panel + claims API-equivalent metered cost, which for a local model is genuinely zero; it does not claim + total cost of ownership, and §4's label says "free" in that specific sense. + +## References + +- Amends [ADR-0009](0009-usage-scorecard-local-transcript-analytics.md) §3, §8; extends + [ADR-0010](0010-provider-mediated-quota-reads.md)'s provider-mediated read pattern to a third + (loopback, credential-free) provider; stays inside [ADR-0005](0005-dashboard-in-page-routing-reveal.md) + / [ADR-0007](0007-maintainer-admin-local-telemetry.md)'s offline `dashboard` contract. +- Ollama — [Anthropic compatibility](https://docs.ollama.com/api/anthropic-compatibility) + ([source `.mdx`](https://github.com/ollama/ollama/blob/main/docs/api/anthropic-compatibility.mdx)): + `/v1/messages`, `usage.{input_tokens, output_tokens}`, unsupported list (prompt caching, token + counting, streaming errors), `ollama cp` aliasing, "token counts are approximations". +- Ollama — [Claude Code integration](https://docs.ollama.com/integrations/claude-code) and + [blog: Claude Code with Anthropic API compatibility](https://ollama.com/blog/claude) + (v0.14.0+; `ANTHROPIC_BASE_URL`, `ANTHROPIC_AUTH_TOKEN`, `ollama launch claude`). +- Ollama — [Codex CLI integration](https://docs.ollama.com/integrations/codex) + (`~/.codex/ollama-launch.config.toml`, `base_url = "http://localhost:11434/v1/"`, + `wire_api = "responses"`, 64k+ context recommendation). +- Ollama — [List models `/api/tags`](https://docs.ollama.com/api/tags), + [List running models `/api/ps`](https://docs.ollama.com/api/ps), + [`docs/api.md`](https://github.com/ollama/ollama/blob/main/docs/api.md) (`/api/show` `model_info`, + `capabilities`; `/api/chat` `prompt_eval_count` / `eval_count`). +- Implementation grounding: [DeepWiki — Anthropic Compatibility Layer](https://deepwiki.com/ollama/ollama/3.5-anthropic-compatibility-layer) + (`anthropic/anthropic.go` `FromMessagesRequest()` / `ToMessagesResponse()`, + `middleware/anthropic.go`; `model` echoed from request; `input_tokens ← prompt_eval_count`, + `output_tokens ← eval_count`). +- Upstream token-accounting instability: [ollama#2068](https://github.com/ollama/ollama/issues/2068) + (`prompt_eval_count` disappears on repeated identical requests), + [ollama#3427](https://github.com/ollama/ollama/issues/3427) (`prompt_eval_count` broken). +- Machine measurements (2026-07-27): `ollama` 0.32.1, daemon down, 132 GB / 14 tags in + `~/.ollama/models/manifests`; usage index schema 6, 754 files, 7 distinct priced model ids, zero + `FALLBACK_PRICE` hits. diff --git a/docs/adr/README.md b/docs/adr/README.md index 0b21861d..a3439863 100644 --- a/docs/adr/README.md +++ b/docs/adr/README.md @@ -18,6 +18,8 @@ Consequences**, and cites the grounded source it rests on where relevant. | [0007](0007-maintainer-admin-local-telemetry.md) | Maintainer admin: a loopback telemetry page with deliberate egress | Accepted | | [0008](0008-guidance-target-scope-split.md) | Machine-scoped guidance blocks land in machine files, not a repo's AGENTS.md | Accepted | | [0009](0009-usage-scorecard-local-transcript-analytics.md) | Usage scorecard: local transcript analytics with graded evidence | Accepted | +| [0010](0010-provider-mediated-quota-reads.md) | Provider-mediated quota reads (the only honest denominators) | Accepted | +| [0011](0011-local-model-provenance-zero-cost-and-transcript-fidelity.md) | Local models: provenance out-of-band, $0 per model, stated transcript fidelity | Proposed | Theme: ADRs **0001–0006** define **dual-host LLM routing and leadership** — how `ak` lets ruflo route each development activity (architecture, implementation, testing, review, …) to the right host (Claude @@ -30,5 +32,13 @@ and folds the two duplicated target lists into one shared `guidanceTargets` help usage scorecard — local transcript analytics as a dashboard tab (no egress, so it stays inside 0005's offline contract rather than joining 0007's egressing admin), with an incremental index over a multi-GB corpus and an explicit evidence grading rule: findings claim a dollar figure only when they -can compute one, and capability claims carry citations or are not made. See also +can compute one, and capability claims carry citations or are not made. **0010** supplies the one +thing 0009 refused to invent — plan-utilisation denominators — by reading each vendor's own reported +percentages through supported channels (Claude Code's statusLine push, Codex's `app-server` RPC), +with provenance and freshness attached. **0011** extends the same discipline to the other end of the +price axis: a session served by a **local** model (`ollama launch claude` / `codex`) reports no cache +accounting, approximate token counts, and a model id the vendor documents how to forge — so +provenance is established out-of-band via a loopback catalogue read, local models are priced at an +exact `$0` **per model** rather than in one bucket, unrecognised models become `unpriced` instead of +silently fallback-priced, and each local session states what its provider could not report. See also `docs/PROVIDERS.md`. From 7e47ef9906b3a5357c24a838376648cb1b505eec Mon Sep 17 00:00:00 2001 From: Chris Phillipson Date: Tue, 28 Jul 2026 00:26:57 -0700 Subject: [PATCH 2/8] feat: reimagine live sessions dashboard --- MAINTAINER.md | 4 +- README.md | 4 +- docs/LIVE-SESSIONS.md | 249 ++ docs/TRANSCRIPTS.md | 48 +- docs/TROUBLESHOOTING.md | 1 + docs/USAGE-SCORECARD-METRICS.md | 46 +- .../0005-dashboard-in-page-routing-reveal.md | 8 +- .../0007-maintainer-admin-local-telemetry.md | 13 +- ...ge-scorecard-local-transcript-analytics.md | 3 + docs/adr/0012-live-sessions-observability.md | 454 +++ docs/adr/README.md | 10 +- docs/ddd/live-sessions.md | 669 ++++ package.json | 2 +- src/commands/x/dashboard.mjs | 50 +- src/lib/admin-server.mjs | 197 +- src/lib/admin-styles.mjs | 189 ++ src/lib/admin-theme.mjs | 26 + src/lib/admin-view.mjs | 4 +- src/lib/codex-state.mjs | 17 +- src/lib/dashboard-server.mjs | 2689 +++-------------- src/lib/dashboard/client.mjs | 1166 +++++++ src/lib/dashboard/live-view.mjs | 3 + src/lib/dashboard/live/client.mjs | 117 + src/lib/dashboard/live/styles.mjs | 20 + src/lib/dashboard/live/template.mjs | 125 + src/lib/dashboard/page.mjs | 264 ++ src/lib/dashboard/request-security.mjs | 37 + src/lib/dashboard/session-security.mjs | 52 + src/lib/dashboard/styles.mjs | 699 +++++ src/lib/live/claude-adapter.mjs | 65 + src/lib/live/codex-adapter.mjs | 151 + src/lib/live/event-schema.mjs | 96 + src/lib/live/index.mjs | 18 + src/lib/live/jsonl-tailer.mjs | 96 + src/lib/live/live-sessions-service.mjs | 288 ++ src/lib/live/project-label.mjs | 22 + src/lib/live/projection.mjs | 271 ++ src/lib/live/replay-stream.mjs | 47 + src/lib/live/service.mjs | 3 + src/lib/live/structured-adapter.mjs | 63 + src/lib/live/tool-classify.mjs | 13 + src/lib/live/transcript-adapter.mjs | 198 ++ src/lib/live/transcript-streams.mjs | 384 +++ tests/admin.test.cjs | 24 + tests/dashboard.test.cjs | 456 +++ tests/kit/codex-state.test.mjs | 17 +- tests/kit/dashboard-live-source.test.mjs | 23 + tests/kit/live-adapters.test.mjs | 139 + tests/kit/live-core.test.mjs | 404 +++ tests/kit/live-qe-contract.test.mjs | 89 + tests/kit/live-service.test.mjs | 214 ++ tests/kit/live-tailer.test.mjs | 53 + tests/kit/live-transcript.test.mjs | 286 ++ tests/ui/dashboard-ui.mjs | 451 ++- 54 files changed, 8569 insertions(+), 2468 deletions(-) create mode 100644 docs/LIVE-SESSIONS.md create mode 100644 docs/adr/0012-live-sessions-observability.md create mode 100644 docs/ddd/live-sessions.md create mode 100644 src/lib/admin-styles.mjs create mode 100644 src/lib/admin-theme.mjs create mode 100644 src/lib/dashboard/client.mjs create mode 100644 src/lib/dashboard/live-view.mjs create mode 100644 src/lib/dashboard/live/client.mjs create mode 100644 src/lib/dashboard/live/styles.mjs create mode 100644 src/lib/dashboard/live/template.mjs create mode 100644 src/lib/dashboard/page.mjs create mode 100644 src/lib/dashboard/request-security.mjs create mode 100644 src/lib/dashboard/session-security.mjs create mode 100644 src/lib/dashboard/styles.mjs create mode 100644 src/lib/live/claude-adapter.mjs create mode 100644 src/lib/live/codex-adapter.mjs create mode 100644 src/lib/live/event-schema.mjs create mode 100644 src/lib/live/index.mjs create mode 100644 src/lib/live/jsonl-tailer.mjs create mode 100644 src/lib/live/live-sessions-service.mjs create mode 100644 src/lib/live/project-label.mjs create mode 100644 src/lib/live/projection.mjs create mode 100644 src/lib/live/replay-stream.mjs create mode 100644 src/lib/live/service.mjs create mode 100644 src/lib/live/structured-adapter.mjs create mode 100644 src/lib/live/tool-classify.mjs create mode 100644 src/lib/live/transcript-adapter.mjs create mode 100644 src/lib/live/transcript-streams.mjs create mode 100644 tests/kit/dashboard-live-source.test.mjs create mode 100644 tests/kit/live-adapters.test.mjs create mode 100644 tests/kit/live-core.test.mjs create mode 100644 tests/kit/live-qe-contract.test.mjs create mode 100644 tests/kit/live-service.test.mjs create mode 100644 tests/kit/live-tailer.test.mjs create mode 100644 tests/kit/live-transcript.test.mjs diff --git a/MAINTAINER.md b/MAINTAINER.md index 92948b37..7f3cc986 100644 --- a/MAINTAINER.md +++ b/MAINTAINER.md @@ -18,7 +18,7 @@ release this kit. User-facing docs live in [README.md](README.md); this file is | Module system | **ESM only** (`.mjs`); `.cjs` reserved for the statusline template + its test | `"type": "module"` in `package.json` | | Runtime deps | **Zero** | Uses `node:` builtins only — `node:sqlite`, `node:test`, `node:util` (`parseArgs`), `node:child_process`, `node:fs`. Keeps installs instant and supply-chain surface tiny | | Node | **≥ 22** (`engines.node`) | `node:sqlite` + `node --test` need it. CI matrix covers 22 / 24 / 26 | -| Package manager (dev) | **pnpm** pinned via `packageManager: pnpm@11.13.0` | CI uses `pnpm/action-setup` which reads that field — don't drift it casually | +| Package manager (dev) | **pnpm** pinned via `packageManager: pnpm@11.17.0` | CI uses `pnpm/action-setup` which reads that field — don't drift it casually | | Package manager (target) | **npm** | The kit heals `npm root -g` trees; `lib/heal.mjs` shells `npm install -g` | | Version source of truth | `package.json` `version` **only** | `bin/agentic-kit.mjs --version` reads it at runtime; no version string is duplicated anywhere else in source | @@ -55,6 +55,8 @@ src/ health-history.mjs # regression ring appended by sync, read by status dashboard-server.mjs # read-only localhost dashboard (shells `ak status --json`) admin-server.mjs # maintainer admin: loopback server, per-session token auth, page assembly (ADR-0007) + admin-styles.mjs # admin presentation: dashboard-aligned dark/light design tokens + admin-theme.mjs # embedded theme controller; shares the dashboard preference key admin-collect.mjs # admin's server-side GitHub/npm fan-out → typed payload (injectable fetchers) admin-model.mjs # PURE admin number model — imports nothing; embedded in the page AND node-tested admin-view.mjs # admin browser controller (embedded into the page; not node-imported) diff --git a/README.md b/README.md index fc04d3ff..eeecac88 100644 --- a/README.md +++ b/README.md @@ -69,8 +69,8 @@ What the verbs cover: | **setup** | Installs/updates ruflo + agentic-qe + the **agentdb** CLI globally (handling npm ≥11.17's `allow-scripts` so natives build; agentdb is pinned to ruflo's bundled version so the shared learning store stays coherent), installs the **RuvNet Brain** (an offline knowledge base over the rUv stack, powering the `search_ruvnet` MCP — a ~2 GB one-time download, prompted; skip with `--no-ruvnet-brain`), deploys the token-audit skill, merges the managed guidance blocks into the machine-wide guidance files (`~/.claude/CLAUDE.md`, plus `~/.codex/AGENTS.md` on codex machines), offers one-time MCP registration (user scope, with a tool-family picker), and — inside a repo — initializes the project: sanitized `ruflo init`, absolute memory-path pin, a **verified** store→disk write, statusline footer, and a background daemon with **local-only ($0) workers** (token-spending AI workers stay opt-in behind upstream's machine-wide budget). Project scope triggers on a `.git` directory in the current folder; without one it's skipped with a note. `--project` forces it anyway (e.g. a not-yet-`git init`-ed folder), `--minimal` skips it, `--yes` accepts all prompts (non-interactive), `--no-aqe` / `--no-ruvnet-brain` / `--no-security` disable those subsystems, and `--reconfigure` re-offers MCP registration. `--codex` enables + installs the Codex host during setup (dual-mode; both hosts then run at once), and `--primary-host claude\|codex` picks which host leads (codex implies `--codex`). | | **status** | Per-subsystem ✓/⚠/✗ (versions, the kit's own version, **ruvnet-brain** (present + release drift, or "not installed"), natives (agentdb copies **and** ruflo's own memory runtime — the one `npx ruflo memory` loads — load-tested for a native better-sqlite3, not just the agentdb dirs), **memory-pin** (warns when `CLAUDE_FLOW_DB_PATH` points off the live DB), security, learning, aqe/RVF, **agentdb** (CLI present + coherent with ruflo's bundled version, or a store-skew warning), MCP, **hosts** (claude/codex — version + install method + **auth mode** (subscription $0 vs metered api-key), with the **primary** host marked and a *fail* when the primary host is absent), **providers** (host wiring + aqe fallback chain, or "drifted"/claude-only default), **routing** (per-activity Claude/Codex host+model policy, when dual-host — with drift vs the on-disk `agentOverrides`), daemons, guidance-file blocks (`~/.claude/CLAUDE.md`, project `AGENTS.md`, and `~/.codex/AGENTS.md` on codex machines), statusline), each drift row naming what `sync` would do about it — plus a **health-history** line that flags regressions since the last sync (learning shrank, native slots dropped, drift/security backslid). | | **sync** | The one convergence verb: upgrades first when a new release exists, then re-heals everything an upgrade wipes, then re-checks and reports. Included in that heal: it **installs any enabled frontier host** (claude/codex) that's entirely absent — never touching an external (mise/brew/native) install — and **re-applies provider wiring** (the `ENABLE_*` host env, the aqe fallback chain, and ruflo API providers) whenever it has drifted — and, on a dual-host project, **seeds/heals the per-activity routing policy** (materializing it into agentic-qe's `agentOverrides`, e.g. after an aqe upgrade first makes it eligible). It also **installs/repins the standalone `agentdb` CLI** to ruflo's bundled version (keeping the shared cognitive store coherent) and appends a **health-history snapshot** so `status` can flag regressions across syncs. It also **re-runs the RuvNet Brain installer** to pull the latest release when the on-disk KB has drifted (or installs it if absent, when enabled). It also **self-updates the kit**: when a newer `@pacphi/agentic-kit` exists it installs it as the *last* step (the new code applies from the next `ak` run, never mid-sync). Prerelease installs (`4.0.0-alpha.*`) track the `next` npm dist-tag as well as `latest`, so alphas see their successors; stable installs only ever follow `latest`. `--no-upgrade` skips the self-update along with the package upgrades. | -| **dashboard** | Opens a read-only local web dashboard (`127.0.0.1:7431`, localhost-only, never detaches) that renders the same subsystem view as `ak status` in an Apple-style five-tab layout — **Overview · Hosts & Routing · Providers · Runtime · Intelligence** — with count badges on any tab holding a failing/warning subsystem. Problems never hide behind a tab: Overview aggregates every attention card, a quiet update notice, and a jump-to status map of all subsystems; Providers shows the **models in play** (distinct host+model pairs from your routing policy); Hosts & Routing carries the **per-activity routing matrix** (vendor-coded Claude/Codex host + model per activity) when dual-host routing is configured; Intelligence keeps the learning-over-time strip. Fully self-contained and offline (no external fetches, nothing leaves your machine). **Auto-opens your browser** (`--no-open` to just print the URL for headless/SSH); `--port N` to change the port; tabs deep-link (`#providers`) and persist. Stop with Ctrl-C. (Also available as `ak x dashboard`.) | -| **admin** | Opens the **maintainer admin** (`127.0.0.1:7432`, localhost-only, foreground) — the project-telemetry sibling of `dashboard`: unique repo visitors and cloners (GitHub traffic API, needs a push-access token via `GITHUB_TOKEN`/`GH_TOKEN`/`gh auth token` — panels degrade honestly without one), npm download momentum (last 7d vs prior 7d, sparklines), release pulls, a **"since you last looked"** delta strip over a local baseline, open issues/PRs from others (oldest first), and external humans ranked by recency (bots excluded). Access is gated by a **per-session token** carried in the URL fragment and sent header-only; the page makes **zero external fetches** (the server proxies GitHub/npm; your credential never reaches the page or the payload — ADR-0007). Where `dashboard` is offline-first, `admin` does deliberate GitHub/npm egress — that contract split is why they're siblings, not tabs. `--port N`, `--no-open`; Ctrl-C stops. (Also available as `ak x admin`.) | +| **dashboard** | Opens a read-only local web dashboard (`127.0.0.1:7431`, localhost-only, never detaches) with seven tabs: **Overview · Hosts & Routing · Providers · Runtime · Intelligence · Usage · Live**. The first five render `ak status` health and routing; Usage indexes local Claude/Codex transcripts on demand. Live groups work by project, then provider-branded root sessions with nested agent/worker threads, and pairs an interactive agent/tool execution canvas with a rich, server-masked transcript stream. Active sessions can be followed live or reviewed with synchronized play/pause/seek; completed sessions remain available for bounded playback. Live contains no chat or control plane. Ruflo, agentic-qe, and dual-run stores are not auto-discovered; register each trusted structured JSONL file with repeatable `--live-source 'surface=path'` (`surface` is `ruflo`, `aqe`, or `dual-run`). The page is self-contained and offline-first (no internet fetches; local files and loopback subprocesses/endpoints only). See [Live Sessions](docs/LIVE-SESSIONS.md) for coverage, syntax, and privacy limits. **Auto-opens your browser** (`--no-open` for headless/SSH); `--port N` changes the port; tabs deep-link (`#live`) and persist. Stop with Ctrl-C. (Also available as `ak x dashboard`.) | +| **admin** | Opens the **maintainer admin** (`127.0.0.1:7432`, localhost-only, foreground) — the project-telemetry sibling of `dashboard`, with the same dark/light visual theme and persisted theme preference: unique repo visitors and cloners (GitHub traffic API, needs a push-access token via `GITHUB_TOKEN`/`GH_TOKEN`/`gh auth token` — panels degrade honestly without one), npm download momentum (last 7d vs prior 7d, sparklines), release pulls, a **"since you last looked"** delta strip over a local baseline, open issues/PRs from others (oldest first), and external humans ranked by recency (bots excluded). Access is gated by a **per-session token** carried in the URL fragment and sent header-only; the page makes **zero external fetches** (the server proxies GitHub/npm; your credential never reaches the page or the payload — ADR-0007). Where `dashboard` is offline-first, `admin` does deliberate GitHub/npm egress — that contract split is why they're siblings, not tabs. `--port N`, `--no-open`; Ctrl-C stops. (Also available as `ak x admin`.) | | **dual** | Runs a **Claude + Codex collaboration swarm** using your per-activity routing policy: `ak dual run