From 948d13a181e0abd53e8fe2d7896618fdee498de2 Mon Sep 17 00:00:00 2001 From: Chris Phillipson Date: Thu, 6 Aug 2026 01:27:53 -0700 Subject: [PATCH] feat(dashboard): usage source-health parity, host-card cleanup, and opt-in runtime debug MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit ADR-0023's sourceHealth field covered only the two secondary/corrective local sources (OpenCode's SQLite store, Codex's thread ledger) and never the primary Claude/Codex transcript roots themselves — an unreadable or missing ~/.claude/projects or ~/.codex/sessions still silently read as an ordinary empty result, the exact failure class ADR-0023 exists to close, just left open on the two sources every installation actually depends on. Verified the fix is grounded, not invented: Anthropic documents the Claude transcript path directly (Data usage page, hooks' transcript_path field); Codex's rollout directory is real and load-bearing for `codex resume`; the Codex ledger and OpenCode's db are real but undocumented internals, confirmed only via their own upstream bug trackers. Also traced and fixed a related Observability gap: runtime process discovery had no way to explain "why didn't this controller show up" short of reading source, and the Usage "by host" panel had drifted from a two-host binary (claude vs codex) to three supported hosts without adjusting its layout or labeling. * Usage source-health parity (ADR-0023 §7) - usage-index.mjs: new rootHealth() reports ok/absent/degraded for the Claude and Codex transcript roots, with fs error codes (ENOENT, EACCES, ENOTDIR) as the bounded reason — same vocabulary already used for OpenCode/codexLedger. - sourceHealth now carries all four fields: claude, codex, opencode, codexLedger. - Dashboard renders by HOST, not by field: three chips for three hosts. Codex's chip folds its transcript-root and thread-ledger statuses together (worse status leads, both sub-statuses stay visible in the chip's detail text) instead of appearing as a fourth, confusingly "duplicate" Codex entry. - Tests: root-health ok/absent/degraded (including an unreadable-root ENOTDIR case) in usage-index.test.mjs; updated dashboard.test.cjs fixtures/assertions for the new grouping. * Usage "by host" panel cleanup - All three hosts (claude/codex/opencode) always render, grayed out via the existing .idle styling when a host has no sessions in the window — never hidden based on whether `ak setup --` was ever run. The scorecard reflects observed transcript evidence, not inferred setup state. - Cards now fit one row (.pcard min-width 190px -> 130px) instead of wrapping OpenCode onto its own row. - Replaced the static, now-inaccurate "claude vs codex" label with a live "N active of 3" count. - OpenCode gets its own dot color (purple) instead of silently sharing Claude's orange. * Opt-in runtime discovery debug flag - process-sessions.mjs: AK_RUNTIME_DEBUG=1 traces the controller-discovery pipeline stage by stage (survey row count, per-PID host classification, nested-child exclusions, cwd resolution, final result count) to $XDG_STATE_HOME/agentic-kit/runtime-debug.log. - Mirrors the existing AK_STATUSLINE_DEBUG contract: off by default, owner-only 0600, bounded/reset at 64 KiB, diagnostic failures can never break discovery. Narrower redaction than the statusline diagnostic by necessity: raw argv/command strings are still never logged, but cwd paths are, since resolving "why didn't project X show up" is the flag's entire purpose and a local directory path isn't a secret. - AK_RUNTIME_DEBUG_FILE overrides the log path (test/operator seam). - Tests: opt-in gating, stage coverage, argv redaction, owner-only mode, and unwritable-sink safety in live-process-sessions.test.mjs. * Docs - ADR-0023: updated §7 (host-grouped chip rationale, grounding citations) and §5 (AK_RUNTIME_DEBUG), update note, Consequences, References. - USAGE-SCORECARD-METRICS.md, TROUBLESHOOTING.md, adr/README.md synced to the four-field/three-chip model and the new debug flag. - TRANSCRIPTS.md / USAGE-SCORECARD-METRICS.md: re-anchored file:line citations that drifted from the usage-index.mjs edits (verified against tests/kit/doc-citations.test.mjs). --- docs/TRANSCRIPTS.md | 10 +-- docs/TROUBLESHOOTING.md | 4 +- docs/USAGE-SCORECARD-METRICS.md | 31 +++++--- ...sed-operations-and-explicit-degradation.md | 56 ++++++++++++-- docs/adr/README.md | 6 +- src/lib/dashboard/client.mjs | 55 +++++++++++--- src/lib/dashboard/page.mjs | 2 +- src/lib/dashboard/styles.mjs | 3 +- src/lib/live/process-sessions.mjs | 55 ++++++++++++-- src/lib/usage-index.mjs | 28 ++++++- tests/dashboard.test.cjs | 4 +- tests/kit/live-process-sessions.test.mjs | 73 +++++++++++++++++++ tests/kit/usage-index.test.mjs | 27 +++++++ 13 files changed, 311 insertions(+), 43 deletions(-) diff --git a/docs/TRANSCRIPTS.md b/docs/TRANSCRIPTS.md index 04340a55..ce27c9c6 100644 --- a/docs/TRANSCRIPTS.md +++ b/docs/TRANSCRIPTS.md @@ -175,15 +175,15 @@ and image-only pastes get the right kind" (the two edges). ## 4. The `readSession` pipeline — how one session becomes a payload -`readSession(id, opts)` (`usage-index.mjs:1271-1359`) is the only way +`readSession(id, opts)` (`usage-index.mjs:1297-1385`) is the only way transcript content leaves the module, and every step is a gate: ### 4.1 Locate, contain, bound 1. **Id grammar before any filesystem access** — `VALID_ID` (`/^[A-Za-z0-9._-]{1,128}$/`, `usage-index.mjs:83`) rejects traversal - shapes with `ERR_INVALID_SESSION_ID` (`usage-index.mjs:1303-1307`). -2. **Locate by id** across both roots (`locate`, `usage-index.mjs:1313`), + shapes with `ERR_INVALID_SESSION_ID` (`usage-index.mjs:1329-1333`). +2. **Locate by id** across both roots (`locate`, `usage-index.mjs:1339`), consulting the scan cache when present but never requiring it — `readSession` works with no prior `buildIndex`. 3. **Realpath containment** (`usage-index.mjs:1335-1349`) — the resolved file @@ -198,7 +198,7 @@ transcript content leaves the module, and every step is a gate: ### 4.2 Parse and price The file is parsed with `withTurns: true` by the provider's parser -(`usage-index.mjs:1404-1411`), and `meta` is assembled +(`usage-index.mjs:1430-1437`), and `meta` is assembled (`usage-index.mjs:1414-1442`) with the same fields the Sessions view rows carry — `prompts`, `responses`, `exceptions`, `sidechain`, `threadSource`, `models`, `tools`, `skill`/`plugin`, worktree — plus a `cost` priced from the @@ -211,7 +211,7 @@ Every turn body is passed through `maskSecrets` (`usage-index.mjs:196` — the 23 secret shapes) **server-side, before serialization**, then length-capped at `MAX_TURN_CHARS` (40,000, `usage-index.mjs:77`) with the marker appended -(`usage-index.mjs:1451-1461`). Two invariants: +(`usage-index.mjs:1477-1487`). Two invariants: - **Presence is the signal.** `truncated`/`originalChars` are emitted only when the slice fired, so a complete turn cannot be misread as abridged. diff --git a/docs/TROUBLESHOOTING.md b/docs/TROUBLESHOOTING.md index ac855379..e7c76189 100644 --- a/docs/TROUBLESHOOTING.md +++ b/docs/TROUBLESHOOTING.md @@ -45,8 +45,8 @@ ak sync # apply it | Observability is empty or has no ruflo/AQE nodes | Live mode tails Claude/Codex records by default, while ruflo/AQE stores are not auto-discovered | open Observability before producing activity; switch to History for retained sessions; register a trusted JSONL file with repeatable `--live-source 'surface=path'`; see [Observability](OBSERVABILITY.md) | | `status` shows `ruvnet-brain … not installed` | The RuvNet Brain (offline KB + `search_ruvnet` MCP) isn't on disk | `ak sync` (or `ak setup`) runs the installer; `npx ruvnet-brain --doctor` health-checks it | | A heal says `degraded` while the tool is still usable | The native repair failed and a fallback or older artifact remains available; exit status is authoritative | Use the reported repair command/error. The operation will not render green or advance a version stamp until a later repair exits successfully | -| Usage/observability suddenly shows no OpenCode or Codex-ledger data | The SQLite source can be absent, busy, corrupt, or query-incompatible; these are no longer collapsed into an ordinary empty result | Inspect the local-source chips at the top of the dashboard Usage area (or `sourceHealth` in usage-index JSON). A degraded OpenCode scan retains in-window last-good cached sessions; repair the named source before treating zero as observed truth | -| Observability does not show a live host process | Runtime discovery uses the numeric UID running the dashboard and is macOS/Linux-only; `sudo`, a service account, Windows, a private container PID namespace, missing `ps`/`lsof`, or restricted `/proc` changes what is visible | Run `ak dashboard` as the same ordinary OS account as the host CLI. Do not use `sudo`; use retained History on Windows and inspect OS/container process permissions when runtime presence is degraded | +| Usage suddenly shows no data for one host, or a lower total than expected | Any of the four local sources (Claude/Codex transcript roots, OpenCode's SQLite store, the Codex thread ledger) can go absent, busy, corrupt, or query-incompatible; none of these are collapsed into an ordinary empty result | Inspect the local-source chips at the top of the dashboard Usage area (or `sourceHealth` in usage-index JSON) — one chip per host; the Codex chip folds its transcript-root and thread-ledger statuses together (worse status leads, both shown in its detail text). A degraded OpenCode scan retains in-window last-good cached sessions; repair the named source before treating zero as observed truth | +| Observability does not show a live host process | Runtime discovery uses the numeric UID running the dashboard and is macOS/Linux-only; `sudo`, a service account, Windows, a private container PID namespace, missing `ps`/`lsof`, or restricted `/proc` changes what is visible | Run `ak dashboard` as the same ordinary OS account as the host CLI. Do not use `sudo`; use retained History on Windows and inspect OS/container process permissions when runtime presence is degraded. If the UID matches and none of the above applies, set `AK_RUNTIME_DEBUG=1` for one reproduction — stage-level evidence (survey row count, host classification per PID, nested-child exclusions, cwd resolution) goes to `$XDG_STATE_HOME/agentic-kit/runtime-debug.log` (mode 0600, bounded at 64 KiB; `AK_RUNTIME_DEBUG_FILE` to redirect it), then unset debug | | Don't want the RuvNet Brain (the ~2 GB KB download) | It's on by default | `ak setup --no-ruvnet-brain`, or set `ruvnetBrain: false` in `~/.config/agentic-kit/kit.json` | | Don't want the security surface managed | Also on by default | `ak setup --no-security` (persists `security:false`; status shows an info row and sync stops healing it) | | RuvNet Brain KB lives somewhere non-default | The installer + ak honor `$RUVNET_BRAIN_KB` (default `~/.cache/ruvnet-brain/kb`) | export `RUVNET_BRAIN_KB` so detection points at your KB | diff --git a/docs/USAGE-SCORECARD-METRICS.md b/docs/USAGE-SCORECARD-METRICS.md index 5dbe3e9c..a8acf068 100644 --- a/docs/USAGE-SCORECARD-METRICS.md +++ b/docs/USAGE-SCORECARD-METRICS.md @@ -72,14 +72,21 @@ Nothing in this transcript pipeline calls a provider API or a billing endpoint; metric is ever a copy of an actual invoice.** That is the whole reason every transcript-derived dollar figure is labelled "API-equivalent." -The built index also exposes `sourceHealth` for the OpenCode SQLite store and -Codex thread ledger. Each source is `ok`, `absent`, `degraded`, or `not-read`, -with a bounded reason such as `busy`, `corrupt`, `query`, or `schema`. A +The built index also exposes `sourceHealth` for all four local sources: the +Claude and Codex transcript roots themselves (`claude`, `codex`), plus the two +secondary/corrective reads layered on top of them (`opencode`'s SQLite store, +`codexLedger`'s thread-attribution ledger). Each source is `ok`, `absent`, +`degraded`, or `not-read`, with a bounded reason such as an fs error code +(`ENOENT`, `EACCES`, `ENOTDIR`), `busy`, `corrupt`, `query`, or `schema`. A degraded OpenCode read retains in-window last-good cached sessions rather than turning an unreadable database into an observed zero. Source health is diagnostic evidence; it is not added to token or cost totals. The dashboard -renders these states as local-source chips above every Usage view so a degraded, -absent, or deliberately unread source cannot be mistaken for healthy empty data. +renders these states as local-source chips above every Usage view, one per +HOST rather than one per field — `codex` and `codexLedger` are both Codex-only +evidence, so they fold into a single "Codex" chip carrying both sub-statuses — +so a degraded, absent, or deliberately unread source cannot be mistaken for +healthy empty data. See [ADR-0023 §7](adr/0023-fail-closed-operations-and-explicit-degradation.md) +for why the four fields are tracked to different degrees of external documentation. The current persisted field named `provider` identifies which host transcript parser produced a session row; it is not sufficient evidence of the inference provider. The Proposed model in @@ -114,9 +121,9 @@ responses = Σ over included sessions of session.responses **Source:** - Filter: a parsed record with zero assistant turns is dropped entirely — "no - assistant turn → not a session" (`usage-index.mjs:904`) — and a record whose + assistant turn → not a session" (`usage-index.mjs:923`) — and a record whose last activity falls outside the requested window is dropped too - (`usage-index.mjs:905`). + (`usage-index.mjs:924`). - `responses` accumulation: Claude increments per assistant message (`usage-index.mjs:504-509`); Codex increments per `agent_message` event (`usage-index.mjs:651-655`). @@ -199,7 +206,7 @@ already in effect on the given day, comparing ISO date strings lexicographically so no `Date` parsing is involved and the module stays clock-free. -`aggregate()` passes each usage row's own `day` (`usage-index.mjs:892`), which +`aggregate()` passes each usage row's own `day` (`usage-index.mjs:911`), which it already has because rows are keyed by `(day, model)`. **This is the whole point:** tokens metered in August must still read as August's rate when the panel is opened in December. Pricing by *today's* date instead would restate a @@ -281,7 +288,7 @@ tokens = input + output + cacheRead + cacheWrite (summed across all rows in wi ``` **Source:** `t.tokens` from `totals`, accumulated per row at -`usage-index.mjs:916` (`rowTokens = row.input + row.output + row.cacheRead + +`usage-index.mjs:935` (`rowTokens = row.input + row.output + row.cacheRead + row.cacheWrite`) and rolled into `totals.tokens` via `addTo` (`usage-index.mjs:843-852`). Rendered with `fmtTok()` (`dashboard/client.mjs`): `≥1e9` → `"X.XB"`, `≥1e6` → `"X.XM"`, @@ -453,7 +460,7 @@ session that runs from 23:58 local to 00:05 local is billed to the day its *first* row landed on (test: `tests/kit/usage-index.test.mjs:634`, "a session that opens before midnight is counted on its first billed day"). Accumulation: -`byDay[row.day].cost += rowCost` (`usage-index.mjs:827`). Bar height: +`byDay[row.day].cost += rowCost` (`usage-index.mjs:846`). Bar height: `h = maxDay ? max(2, cost/maxDay*100) : 2` (`dashboard/client.mjs`) — every non-empty day gets a visually nonzero bar (floor of 2%), so a very cheap day is never rendered as invisible. @@ -474,7 +481,7 @@ renders "no sessions in window" instead of zeroed figures (`dashboard/client.mjs`). **Formula:** identical aggregation to every other bucket -(`byProvider[s.provider]`, populated via `addTo()`, `usage-index.mjs:870-879`, +(`byProvider[s.provider]`, populated via `addTo()`, `usage-index.mjs:889-898`, called once per session at `usage-index.mjs:1000`), keyed by the literal string `"claude"` or `"codex"` assigned at parse time (`blankSession(id, 'claude')` / `blankSession(id, 'codex')`, @@ -516,7 +523,7 @@ punchcard[dow + "-" + hour] += 1 per assistant/agent_message response, at its **Source:** incremented once per Claude assistant turn (`usage-index.mjs:504-509`, keyed by `punchKey(at)`) and once per Codex `agent_message` (`usage-index.mjs:651-655`), merged into the window-level -`punchcard` object per session (`usage-index.mjs:981`). Cell intensity is +`punchcard` object per session (`usage-index.mjs:1000`). Cell intensity is linear against the single busiest cell in the window: `v = pcMax ? n/pcMax : 0` (`dashboard/client.mjs`) — this is a **relative**, not absolute, scale, so the heatmap's brightest cell is always diff --git a/docs/adr/0023-fail-closed-operations-and-explicit-degradation.md b/docs/adr/0023-fail-closed-operations-and-explicit-degradation.md index 6189def3..ab6b5df9 100644 --- a/docs/adr/0023-fail-closed-operations-and-explicit-degradation.md +++ b/docs/adr/0023-fail-closed-operations-and-explicit-degradation.md @@ -2,10 +2,14 @@ - **Status:** Implemented - **Date:** 2026-08-04 -- **Updated:** 2026-08-04 +- **Updated:** 2026-08-06 - **Update note:** Generalized setup preflight into a required host-adapter trust contract, added Codex registration/OpenCode approval disclosure, documented current-UID installation-mode - boundaries, and surfaced usage-source health in the dashboard UI. + boundaries, and surfaced usage-source health in the dashboard UI. Closed a parity gap §7 left + behind: `sourceHealth` covered only the two secondary/corrective sources (OpenCode's SQLite + store, Codex's thread ledger) and never the primary Claude/Codex transcript roots, so a missing + or unreadable `~/.claude/projects` or `~/.codex/sessions` still silently read as zero — the exact + failure class this ADR exists to close, just left open on the two sources most people depend on. - **Deciders:** agentic-kit maintainers - **Related:** [issue #111](https://github.com/pacphi/agentic-kit/issues/111), [ADR-0008](0008-guidance-target-scope-split.md), @@ -82,6 +86,15 @@ survey omits argv. A second query reads argv only for current-user executables t host controller or Node launcher. CWD lookup then receives only filtered controller PIDs. The public event boundary remains path-redacted as defined by ADR-0012. +`AK_RUNTIME_DEBUG=1` traces this pipeline stage-by-stage (survey row count, which PIDs were classified +as a host, which were dropped as a nested child of another candidate, and per-controller cwd +resolution) to `$XDG_STATE_HOME/agentic-kit/runtime-debug.log`, mirroring the statusline diagnostic's +opt-in/bounded-at-64-KiB/owner-only-0600 contract (`AK_RUNTIME_DEBUG_FILE` overrides the path). It is a +narrower redaction than the statusline diagnostic: raw argv/command strings are still never logged +(a pasted prompt or token could be sitting in one), but cwd paths ARE — resolving "why didn't project +X's controller show up" is the flag's entire purpose, the operator turned it on deliberately, and a +local directory path is not a secret the way a command line can be. + ### 6. Every host setup has a pre-mutation trust boundary Each host adapter must declare whether agentic-kit manages approval grants or leaves the host's @@ -99,11 +112,41 @@ rules survive. OpenCode discloses its user-scope wildcard approvals, MCP registr plugin, and managed host assets. Codex discloses MCP/AQE registrations while explicitly retaining its sandbox and approval policy. -### 7. Usage source degradation is visible in the dashboard +### 7. Usage source degradation is visible in the dashboard, for all four local sources The Usage API's `sourceHealth` field is rendered as persistent local-source chips across Usage views. `ok`, `absent`, `degraded`, and `not-read` remain distinct, and bounded reasons such as -`busy`, `corrupt`, `query`, `schema`, or `sandboxed-roots` are visible without entering raw JSON. +`busy`, `corrupt`, `query`, `schema`, `sandboxed-roots`, or an fs error code (`ENOENT`, `EACCES`, +`ENOTDIR`) are visible without entering raw JSON. + +`sourceHealth` originally covered only the two sources with a *secondary, corrective* read layered +on top of a primary parse — OpenCode's SQLite store and Codex's own thread ledger — because those +are exactly where item 2 above found silent collapse in practice. It did not cover the primary +Claude and Codex transcript roots (`~/.claude/projects`, `~/.codex/sessions`) themselves: `listClaude` +and `listCodex` walk those directories through a `readdirSync` wrapped in a bare `catch { return [] }`, +so a missing root, a permissions error, or any other I/O failure was indistinguishable from "no +sessions in this window" — the same silent-zero failure class item 2 closed for OpenCode/Codex, just +left open on the two sources every installation actually depends on. + +Closing it required checking what's real to check against, not inventing a status: Anthropic +documents `~/.claude/projects//*.jsonl` and its 30-day default retention directly (Claude +Code's Data usage page; the `transcript_path` every hook receives). Codex's `~/.codex/sessions/**/rollout-*.jsonl` +is real and load-bearing for Codex's own `codex resume`, though OpenAI does not publish it as a +formal contract the way Anthropic does. Both are markedly more stable ground than the two sources +already tracked — the Codex thread ledger (`state_N.sqlite`) and OpenCode's `opencode.db` are +undocumented internal storage, confirmed real only by their own upstream bug trackers (e.g. +openai/codex#21750, a corrupt `state_5.sqlite` wedging Codex's own startup) and defensive code +already treating the ledger's generation suffix as "not a stable name." None of the four are +fabricated; `rootHealth()` performs a real `readdirSync` against a real path exactly like the +existing checks, just one level up the trust stack from the two sources already wired. + +`sourceHealth` now reports `claude`, `codex`, `opencode`, and `codexLedger`. The dashboard renders +this by HOST, not by field: three chips for the three supported hosts (Claude, Codex, OpenCode), not +four. `codex` and `codexLedger` are both Codex-only evidence, so they fold into one "Codex" chip — +its status is the worse of the two, and both sub-statuses stay visible in the chip's detail text +(e.g. `Codex: degraded — transcripts: ok · ledger: corrupt`). No evidence is dropped; the API keeps +four independently-diagnosable fields, the UI just groups by the thing the operator actually cares +about (which host needs attention), matching how Claude and OpenCode already render as one chip each. ### 8. Clean-machine proof is isolated at every mutable boundary @@ -124,6 +167,9 @@ setup on `macos-latest` with all global packages and user/project files under `r undisclosed Claude project grants introduced by upstream initializers. - Some formerly best-effort writes now fail. This is deliberate: when ak promises a backup, mutation without one is a correctness failure. +- An unreadable `~/.claude/projects` or `~/.codex/sessions` (permissions, a corrupt filesystem entry, + the path replaced by a non-directory) now renders as a degraded local-source chip instead of a + quietly empty Usage scorecard; all four local sources share one status vocabulary. ## References @@ -132,6 +178,6 @@ setup on `macos-latest` with all global packages and user/project files under `r `src/lib/dashboard/{page,client,styles}.mjs`, `src/commands/{setup,sync}.mjs`, and `src/templates/statusline-footer.cjs`. - Tests: `tests/kit/{clean-machine-setup,heal-natives,sqlite,settings-config,blocks, - live-process-sessions,setup-command,trust-manifest,usage-index-opencode}.test.mjs`, + live-process-sessions,setup-command,trust-manifest,usage-index,usage-index-opencode}.test.mjs`, `tests/dashboard.test.cjs`, and `tests/statusline-segments.test.cjs`. - Clean-machine workflow: `.github/workflows/nightly.yml`. diff --git a/docs/adr/README.md b/docs/adr/README.md index e77cd4a5..86012e02 100644 --- a/docs/adr/README.md +++ b/docs/adr/README.md @@ -121,7 +121,11 @@ interop API. managed fallbacks report degradation, SQLite retains classified failure evidence and last-good usage, promised backups fail closed before atomic replacement, status-line failures gain redacted opt-in diagnostics, process discovery is current-user and argv-minimized, setup discloses and -verifies its project auto-approve manifest, and clean-machine tests isolate every mutable path. +verifies its project auto-approve manifest, and clean-machine tests isolate every mutable path. A +follow-up closed a parity gap in the last item: `sourceHealth` originally covered only the two +secondary/corrective sources (OpenCode's store, Codex's thread ledger), not the primary Claude/Codex +transcript roots — an unreadable `~/.claude/projects` or `~/.codex/sessions` still read as silent +zero. It now reports all four. **0024** gives Overview's Intelligence view real trend data instead of a permanently-empty strip: a new `intel-history.mjs` module reads the neural pattern store, its lifetime learned-pattern diff --git a/src/lib/dashboard/client.mjs b/src/lib/dashboard/client.mjs index e16a06bd..5f34f6e2 100644 --- a/src/lib/dashboard/client.mjs +++ b/src/lib/dashboard/client.mjs @@ -791,15 +791,43 @@ export const JS = ` return out; } + // One chip per HOST, not per sourceHealth field: Codex carries two fields + // (its transcript root, plus the thread ledger that corrects subagent-replay + // token inflation) but is still one host, so the worse of the two drives the + // chip's status and both are visible in its detail text. + var SOURCE_HEALTH_GROUPS=[ + {label:"Claude",title:"~/.claude/projects/**/*.jsonl",parts:[{key:"claude"}]}, + {label:"Codex",title:"~/.codex/sessions/**/rollout-*.jsonl, plus Codex's own thread ledger", + parts:[{key:"codex",sub:"transcripts"},{key:"codexLedger",sub:"ledger"}]}, + {label:"OpenCode",title:"OpenCode's SQLite session store",parts:[{key:"opencode"}]} + ]; + var SOURCE_HEALTH_RANK={degraded:0,"not-read":1,absent:2,ok:3}; // lower sorts first = worse + function sourceHealthRank(status){ + var r=SOURCE_HEALTH_RANK[status]; + return r===undefined?1:r; + } + function renderSourceHealth(health){ var el=document.getElementById("u-source-health"); if(!el)return; - var labels={opencode:"OpenCode",codexLedger:"Codex ledger"},chips=[]; - for(var key in (health||{})){ - var item=health[key]||{},status=String(item.status||"not-read"); - var detail=status+(item.reason?" · "+item.reason:""); - chips.push('' - +esc(labels[key]||key)+": "+esc(detail)+""); + health=health||{}; + var chips=[]; + for(var g=0; g' + +esc(grp.label)+": "+esc(detail)+""); } el.hidden=chips.length===0; el.innerHTML=chips.length?'local sources'+chips.join(""):""; @@ -912,18 +940,27 @@ export const JS = ` +''+esc(x.day.slice(8))+""; }).join(""):'
no days in window.
'; - // Host and inference-provider are independent canonical axes. + // Host and inference-provider are independent canonical axes. All three + // supported hosts always render (idle/grayed-out when a host has no + // sessions in this window) rather than appearing/disappearing based on + // setup state — the scorecard reflects observed transcript evidence, not + // which ak setup --host flags were ever run. var prov=d.byHost||{}; - var order=["claude","codex"]; + var order=["claude","codex","opencode"]; for(k in prov)if(order.indexOf(k)<0)order.push(k); + var PDOT_CLASS={claude:"c",codex:"x",opencode:"o"}; + var activeHosts=0; document.getElementById("u-hosts").innerHTML=order.map(function(name){ var v=prov[name], cost=fld(v,"cost"), sess=fld(v,"sessions"), tok=fld(v,"tokens"); var idle=!sess&&!cost; + if(!idle)activeHosts++; return '
'+esc(name)+"
" + +(PDOT_CLASS[name]||"c")+'">'+esc(name)+"
" +'
'+esc(fmtUsd(cost))+"
" +'
'+(idle?"no sessions in window":esc(fmtNum(sess))+" sessions · "+esc(fmtTok(tok))+" tokens")+"
"; }).join(""); + var hostsNoteEl=document.getElementById("u-hosts-note"); + if(hostsNoteEl)hostsNoteEl.textContent=activeHosts+" active of "+order.length; var segs=[["cache read",t.cacheRead,"var(--warn)"],["cache write",t.cacheWrite,"var(--purple)"], ["output",t.output,"var(--accent)"],["input",t.input,"var(--ok)"]]; diff --git a/src/lib/dashboard/page.mjs b/src/lib/dashboard/page.mjs index aabb2b56..847a2258 100644 --- a/src/lib/dashboard/page.mjs +++ b/src/lib/dashboard/page.mjs @@ -284,7 +284,7 @@ export function renderPage({ name, version }) {
-

by host

claude vs codex
+

by host

diff --git a/src/lib/dashboard/styles.mjs b/src/lib/dashboard/styles.mjs index 8bbd1abe..1ae03ad5 100644 --- a/src/lib/dashboard/styles.mjs +++ b/src/lib/dashboard/styles.mjs @@ -476,11 +476,12 @@ body.gated .band,body.gated .tabbar,body.gated main{display:none} /* provider split */ .psplit{display:flex; gap:10px; flex-wrap:wrap} -.pcard{flex:1; min-width:190px; border:1px solid var(--line); border-radius:var(--r-sm); padding:13px 15px; background:var(--panel-2)} +.pcard{flex:1; min-width:130px; border:1px solid var(--line); border-radius:var(--r-sm); padding:13px 15px; background:var(--panel-2)} .pcard .ph{display:flex; align-items:center; gap:8px; font-size:13px; font-weight:600; margin-bottom:9px} .pdot{width:9px; height:9px; border-radius:50%} .pdot.c{background:var(--warn)} .pdot.x{background:var(--accent)} +.pdot.o{background:var(--purple)} .pcard .pv{font-size:21px; font-weight:700; letter-spacing:-.02em} .pcard .pl{font-size:11.5px; color:var(--ink-dim); margin-top:4px} .pcard.idle{opacity:.55} diff --git a/src/lib/live/process-sessions.mjs b/src/lib/live/process-sessions.mjs index 2f299927..5e1f338f 100644 --- a/src/lib/live/process-sessions.mjs +++ b/src/lib/live/process-sessions.mjs @@ -1,10 +1,40 @@ import { execFile } from 'node:child_process'; import fs from 'node:fs'; +import os from 'node:os'; import path from 'node:path'; import { promisify } from 'node:util'; import { inspectGitWorkspace } from './git-workspace.mjs'; const execFileAsync = promisify(execFile); + +// Off by default (ADR-0023 §5/§4 precedent: AK_STATUSLINE_DEBUG). Set +// AK_RUNTIME_DEBUG=1 to trace why a controller process was or wasn't +// surfaced — which rows the ps survey saw, which were classified as a +// host, which got dropped as a nested child of another candidate, and +// whether cwd resolution found each controller. Unlike the statusline +// diagnostic (which runs on every keystroke for users who may not even +// know the flag exists), this is opt-in and single-purpose, so the cwd +// path IS logged — it's the exact fact being debugged and it's not a +// secret. Raw argv/command strings are still never logged: a pasted +// prompt or token could be sitting in there. +function runtimeDebug(stage, fields = {}) { + if (!process?.env || process.env.AK_RUNTIME_DEBUG !== '1') return; + try { + const root = process.env.XDG_STATE_HOME || path.join(os.homedir(), '.local', 'state'); + const file = process.env.AK_RUNTIME_DEBUG_FILE || path.join(root, 'agentic-kit', 'runtime-debug.log'); + const safeStage = String(stage || 'unknown').replace(/[^a-z0-9._-]/gi, '_').slice(0, 64); + const kv = Object.entries(fields) + .map(([k, v]) => `${k}=${String(v).replace(/[\s\n]+/g, '_').slice(0, 200)}`) + .join(' '); + const line = `${new Date().toISOString()} stage=${safeStage} ${kv}\n`; + fs.mkdirSync(path.dirname(file), { recursive: true, mode: 0o700 }); + let size = 0; + try { size = fs.statSync(file).size; } catch { /* first write */ } + if (size >= 65536) fs.writeFileSync(file, line, { mode: 0o600 }); + else fs.appendFileSync(file, line, { mode: 0o600 }); + try { fs.chmodSync(file, 0o600); } catch { /* best-effort */ } + } catch { /* diagnostics must never break discovery */ } +} const HOST_NAMES = new Map([ ['claude', 'claude'], ['codex', 'codex'], @@ -76,17 +106,23 @@ function rootControllers(rows) { for (const row of rows) { const host = hostFromCommand(row.command, row.executable); if (host) candidates.set(row.pid, { ...row, host }); + else if (row.command) runtimeDebug('classify', { pid: row.pid, exe: row.executable, host: 'none' }); } - return [...candidates.values()].filter((candidate) => { + const roots = [...candidates.values()].filter((candidate) => { const seen = new Set([candidate.pid]); let parent = byPid.get(candidate.ppid); while (parent && !seen.has(parent.pid)) { - if (candidates.has(parent.pid)) return false; + if (candidates.has(parent.pid)) { + runtimeDebug('exclude-nested-child', { pid: candidate.pid, host: candidate.host, parentPid: parent.pid }); + return false; + } seen.add(parent.pid); parent = byPid.get(parent.ppid); } return true; }); + for (const root of roots) runtimeDebug('root-controller', { pid: root.pid, host: root.host }); + return roots; } async function linuxCwds(pids) { @@ -168,6 +204,7 @@ export async function listActiveHostSessions({ env: { ...process.env, LC_ALL: 'C' } }); const output = typeof result === 'string' ? result : result.stdout; rows = parseProcessHeaders(output); + runtimeDebug('survey', { uid, rowCount: rows.length }); if (String(output ?? '').trim() && !rows.length) { throw Object.assign(new Error('runtime process output was not understood'), { code: 'ERR_RUNTIME_PROCESS_FORMAT', @@ -179,6 +216,7 @@ export async function listActiveHostSessions({ const name = executableName(row.executable); return HOST_NAMES.has(name) || name === 'node' || name === 'nodejs'; }); + runtimeDebug('argv-candidates', { count: possible.length, pids: possible.map((r) => r.pid).join(',') }); if (possible.length) { const argsResult = await execFileImpl('ps', [ '-p', possible.map((row) => row.pid).join(','), '-o', 'pid=,args=', @@ -188,7 +226,8 @@ export async function listActiveHostSessions({ rows = rows.map((row) => commands.has(row.pid) ? { ...row, command: commands.get(row.pid) } : row); } - } catch { + } catch (error) { + runtimeDebug('survey-failed', { name: error?.name, code: error?.code }); throw Object.assign(new Error('runtime process survey failed'), { code: 'ERR_RUNTIME_PROCESS_SURVEY', }); @@ -198,17 +237,23 @@ export async function listActiveHostSessions({ const pids = controllers.map((row) => row.pid); const cwds = cwdByPid ?? (platform === 'linux' ? await linuxCwds(pids) : await darwinCwds(pids, execFileImpl)); + for (const pid of pids) runtimeDebug('cwd', { pid, found: cwds.has(pid), cwd: cwds.get(pid) ?? '' }); const workspaceByCwd = new Map(); await Promise.all([...new Set(cwds.values())].map(async (cwd) => { try { workspaceByCwd.set(cwd, await inspectWorkspace(cwd)); } catch { workspaceByCwd.set(cwd, null); } })); - return controllers.flatMap((row) => { + const sessions = controllers.flatMap((row) => { const cwd = cwds.get(row.pid); - if (typeof cwd !== 'string' || !path.isAbsolute(cwd)) return []; + if (typeof cwd !== 'string' || !path.isAbsolute(cwd)) { + runtimeDebug('drop-no-cwd', { pid: row.pid, host: row.host }); + return []; + } const workspace = workspaceByCwd.get(cwd); return workspace ? [{ pid: row.pid, startedAt: row.startedAt, host: row.host, cwd, workspace }] : [{ pid: row.pid, startedAt: row.startedAt, host: row.host, cwd }]; }).sort((left, right) => left.pid - right.pid); + runtimeDebug('result', { sessionCount: sessions.length }); + return sessions; } diff --git a/src/lib/usage-index.mjs b/src/lib/usage-index.mjs index 7098ad29..9594f453 100644 --- a/src/lib/usage-index.mjs +++ b/src/lib/usage-index.mjs @@ -694,6 +694,25 @@ function statSafe(file) { try { return fs.statSync(file); } catch { return null; } } +/** ok/absent/degraded for a primary transcript root (Claude/Codex): distinguishes + * a directory that simply doesn't exist yet (host never used on this machine) + * from one that exists but can't be read (permissions, not-a-directory, I/O) — + * the same vocabulary sourceHealth already uses for the opencode/codexLedger + * sources, so an unreadable root cannot be mistaken for "zero sessions in + * this window". listClaude/listCodex still swallow readdir errors per + * subdirectory (a bad nested entry must not abort the whole scan); this is + * the one root-level check that turns that silence into visible evidence. */ +function rootHealth(dir) { + try { + fs.readdirSync(dir); + return { status: 'ok', reason: null }; + } catch (err) { + return err?.code === 'ENOENT' + ? { status: 'absent', reason: null } + : { status: 'degraded', reason: err?.code || 'io' }; + } +} + function defaultRoots() { return { claude: path.join(claudeDir(), 'projects'), @@ -1136,6 +1155,11 @@ async function scan(o = {}) { const cacheFile = cachePath ?? defaultCachePath(); const cutoff = now - days * DAY_MS; + // Primary transcript roots: read once at root level (cheap — not the + // recursive per-file walk listClaude/listCodex still do below). + const claudeHealth = rootHealth(r.claude); + const codexHealth = rootHealth(r.codex); + const candidates = [...listClaude(r.claude), ...listCodex(r.codex)] .map((e) => ({ ...e, stat: statSafe(e.file) })) .filter((e) => e.stat && e.stat.mtimeMs >= cutoff); @@ -1244,7 +1268,9 @@ async function scan(o = {}) { : { status: observed.error.kind === 'absent' ? 'absent' : 'degraded', reason: observed.error.kind }; } const result = aggregate(applyCodexLedger(records, ledger), { days, now, cutoff, deps }); - result.sourceHealth = { opencode: opencodeHealth, codexLedger: codexLedgerHealth }; + result.sourceHealth = { + claude: claudeHealth, codex: codexHealth, opencode: opencodeHealth, codexLedger: codexLedgerHealth, + }; return result; } diff --git a/tests/dashboard.test.cjs b/tests/dashboard.test.cjs index 47924283..74b84a5b 100644 --- a/tests/dashboard.test.cjs +++ b/tests/dashboard.test.cjs @@ -445,6 +445,8 @@ async function main() { const AGG = { generatedAt: '2026-07-25T00:00:00.000Z', windowDays: 14, pricesAsOf: '2026-07-01', sourceHealth: { + claude: { status: 'ok', reason: null }, + codex: { status: 'ok', reason: null }, opencode: { status: 'degraded', reason: 'busy' }, codexLedger: { status: 'ok', reason: null }, }, @@ -672,7 +674,7 @@ async function main() { contains(r.body, 'function renderSourceHealth'); contains(r.body, 'data-status="'); contains(r.body, 'OpenCode'); - contains(r.body, 'Codex ledger'); + contains(r.body, 'SOURCE_HEALTH_GROUPS'); // Codex + its thread ledger render as one grouped chip contains(r.body, 'provider account analytics'); contains(r.body, 'never merged into transcript totals'); contains(r.body, 'OpenRouter credits'); diff --git a/tests/kit/live-process-sessions.test.mjs b/tests/kit/live-process-sessions.test.mjs index bdb2298f..98392d35 100644 --- a/tests/kit/live-process-sessions.test.mjs +++ b/tests/kit/live-process-sessions.test.mjs @@ -1,5 +1,8 @@ import { test } from 'node:test'; import assert from 'node:assert/strict'; +import fs from 'node:fs'; +import os from 'node:os'; +import path from 'node:path'; import { hostFromCommand, listActiveHostSessions, parseLsofCwds, parseProcessHeaders, parseProcessList, } from '../../src/lib/live/index.mjs'; @@ -115,3 +118,73 @@ test('runtime discovery scopes ps to the current UID and reads argv only for hos assert.equal(calls[0].args.includes('-a'), false, 'never surveys all users'); assert.equal(calls[1].args[1], '100,102', 'ssh/non-host PID never reaches the argv survey'); }); + +test('AK_RUNTIME_DEBUG is opt-in, bounded to known stages, and never leaks raw argv', async () => { + const startedAt = 'Mon Aug 3 12:00:00 2026'; + const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'ak-runtime-debug-')); + const log = path.join(dir, 'runtime-debug.log'); + const execFileImpl = async (command, args) => { + if (args.includes('pid=,ppid=,lstart=,comm=')) { + return { stdout: [ + `100 1 ${startedAt} claude`, + `101 100 ${startedAt} ssh`, + ].join('\n') }; + } + if (args.includes('pid=,args=')) return { stdout: '100 claude --dangerous-secret-token\n' }; + throw new Error(`unexpected command: ${command} ${args.join(' ')}`); + }; + const before = process.env.AK_RUNTIME_DEBUG; + const beforeFile = process.env.AK_RUNTIME_DEBUG_FILE; + try { + delete process.env.AK_RUNTIME_DEBUG; + process.env.AK_RUNTIME_DEBUG_FILE = log; + await listActiveHostSessions({ + platform: 'linux', uid: 501, execFileImpl, + cwdByPid: new Map([[100, '/repos/emailibrium']]), + inspectWorkspace: async () => null, + }); + assert(!fs.existsSync(log), 'debug-off must not write a diagnostic log'); + + process.env.AK_RUNTIME_DEBUG = '1'; + await listActiveHostSessions({ + platform: 'linux', uid: 501, execFileImpl, + cwdByPid: new Map([[100, '/repos/emailibrium']]), + inspectWorkspace: async () => null, + }); + const diagnostic = fs.readFileSync(log, 'utf8'); + assert.match(diagnostic, /stage=survey uid=501 rowCount=2/); + assert.match(diagnostic, /stage=argv-candidates count=1 pids=100/); + assert.match(diagnostic, /stage=root-controller pid=100 host=claude/); + assert.match(diagnostic, /stage=cwd pid=100 found=true cwd=\/repos\/emailibrium/); + assert.match(diagnostic, /stage=result sessionCount=1/); + assert(!diagnostic.includes('--dangerous-secret-token'), 'raw argv must never reach the debug log'); + if (process.platform !== 'win32') { + assert.equal(fs.statSync(log).mode & 0o777, 0o600, 'debug log must be owner-only'); + } + } finally { + if (before === undefined) delete process.env.AK_RUNTIME_DEBUG; else process.env.AK_RUNTIME_DEBUG = before; + if (beforeFile === undefined) delete process.env.AK_RUNTIME_DEBUG_FILE; else process.env.AK_RUNTIME_DEBUG_FILE = beforeFile; + fs.rmSync(dir, { recursive: true, force: true }); + } +}); + +test('an unwritable runtime-debug sink never breaks discovery', async () => { + const startedAt = 'Mon Aug 3 12:00:00 2026'; + const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'ak-runtime-debug-unwritable-')); + const before = process.env.AK_RUNTIME_DEBUG; + const beforeFile = process.env.AK_RUNTIME_DEBUG_FILE; + try { + process.env.AK_RUNTIME_DEBUG = '1'; + process.env.AK_RUNTIME_DEBUG_FILE = dir; // appendFileSync on a directory fails + const processRows = parseProcessList([`100 1 ${startedAt} claude claude`].join('\n')); + const sessions = await listActiveHostSessions({ + platform: 'darwin', processRows, cwdByPid: new Map([[100, '/repos/keel']]), + inspectWorkspace: async () => null, + }); + assert.deepEqual(sessions.map((s) => s.host), ['claude']); + } finally { + if (before === undefined) delete process.env.AK_RUNTIME_DEBUG; else process.env.AK_RUNTIME_DEBUG = before; + if (beforeFile === undefined) delete process.env.AK_RUNTIME_DEBUG_FILE; else process.env.AK_RUNTIME_DEBUG_FILE = beforeFile; + fs.rmSync(dir, { recursive: true, force: true }); + } +}); diff --git a/tests/kit/usage-index.test.mjs b/tests/kit/usage-index.test.mjs index 944b605f..d51ba1f0 100644 --- a/tests/kit/usage-index.test.mjs +++ b/tests/kit/usage-index.test.mjs @@ -769,6 +769,33 @@ test('an empty corpus yields a zeroed Aggregate rather than throwing', async () assert.equal(agg.totals.engagedSeconds, 0); assert.deepEqual(agg.sessions, []); assert.deepEqual(agg.projectTree, []); + // A never-used host (root simply doesn't exist yet) reads as absent, not ok — + // "zero sessions" and "we never found the directory" must stay distinguishable. + assert.deepEqual(agg.sourceHealth.claude, { status: 'absent', reason: null }); + assert.deepEqual(agg.sourceHealth.codex, { status: 'absent', reason: null }); +}); + +test('buildIndex reports ok claude/codex root health when the transcript roots exist', async () => { + _resetForTest(); + const sb = sandbox(); + const agg = await buildIndex(opts(sb)); + assert.deepEqual(agg.sourceHealth.claude, { status: 'ok', reason: null }); + assert.deepEqual(agg.sourceHealth.codex, { status: 'ok', reason: null }); +}); + +test('an unreadable Claude root degrades rather than silently reading as zero sessions', async () => { + _resetForTest(); + const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'ak-usage-unreadable-')); + const claudeRoot = path.join(dir, 'claude-is-a-file'); + fs.writeFileSync(claudeRoot, 'not a directory'); // readdirSync on this throws ENOTDIR, not ENOENT + const agg = await buildIndex({ + days: 14, now: NOW, deps: deps(), + roots: { claude: claudeRoot, codex: path.join(dir, 'codex-nope') }, + cachePath: path.join(dir, 'cache.json'), + }); + assert.equal(agg.sourceHealth.claude.status, 'degraded'); + assert.ok(agg.sourceHealth.claude.reason, 'a degraded root must carry a bounded reason, e.g. ENOTDIR'); + assert.equal(agg.totals.sessions, 0, 'degraded still yields zero sessions here (nothing to preserve) — the point is the status, not the count'); }); // ── cache ───────────────────────────────────────────────────────────────────