Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 5 additions & 5 deletions docs/TRANSCRIPTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -175,15 +175,15 @@ and image-only pastes get the right kind" (the two edges).

## 4. The `readSession` pipeline — how one session becomes a payload

`readSession(id, opts)` (`usage-index.mjs:1271-1359`) is the only way
`readSession(id, opts)` (`usage-index.mjs:1297-1385`) is the only way
transcript content leaves the module, and every step is a gate:

### 4.1 Locate, contain, bound

1. **Id grammar before any filesystem access** — `VALID_ID`
(`/^[A-Za-z0-9._-]{1,128}$/`, `usage-index.mjs:83`) rejects traversal
shapes with `ERR_INVALID_SESSION_ID` (`usage-index.mjs:1303-1307`).
2. **Locate by id** across both roots (`locate`, `usage-index.mjs:1313`),
shapes with `ERR_INVALID_SESSION_ID` (`usage-index.mjs:1329-1333`).
2. **Locate by id** across both roots (`locate`, `usage-index.mjs:1339`),
consulting the scan cache when present but never requiring it —
`readSession` works with no prior `buildIndex`.
3. **Realpath containment** (`usage-index.mjs:1335-1349`) — the resolved file
Expand All @@ -198,7 +198,7 @@ transcript content leaves the module, and every step is a gate:
### 4.2 Parse and price

The file is parsed with `withTurns: true` by the provider's parser
(`usage-index.mjs:1404-1411`), and `meta` is assembled
(`usage-index.mjs:1430-1437`), and `meta` is assembled
(`usage-index.mjs:1414-1442`) with the same fields the Sessions view rows
carry — `prompts`, `responses`, `exceptions`, `sidechain`, `threadSource`,
`models`, `tools`, `skill`/`plugin`, worktree — plus a `cost` priced from the
Expand All @@ -211,7 +211,7 @@ Every turn body is passed through `maskSecrets` (`usage-index.mjs:196` — the
23 secret shapes) **server-side, before
serialization**, then length-capped at `MAX_TURN_CHARS` (40,000,
`usage-index.mjs:77`) with the marker appended
(`usage-index.mjs:1451-1461`). Two invariants:
(`usage-index.mjs:1477-1487`). Two invariants:

- **Presence is the signal.** `truncated`/`originalChars` are emitted only
when the slice fired, so a complete turn cannot be misread as abridged.
Expand Down
4 changes: 2 additions & 2 deletions docs/TROUBLESHOOTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -45,8 +45,8 @@ ak sync # apply it
| Observability is empty or has no ruflo/AQE nodes | Live mode tails Claude/Codex records by default, while ruflo/AQE stores are not auto-discovered | open Observability before producing activity; switch to History for retained sessions; register a trusted JSONL file with repeatable `--live-source 'surface=path'`; see [Observability](OBSERVABILITY.md) |
| `status` shows `ruvnet-brain … not installed` | The RuvNet Brain (offline KB + `search_ruvnet` MCP) isn't on disk | `ak sync` (or `ak setup`) runs the installer; `npx ruvnet-brain --doctor` health-checks it |
| A heal says `degraded` while the tool is still usable | The native repair failed and a fallback or older artifact remains available; exit status is authoritative | Use the reported repair command/error. The operation will not render green or advance a version stamp until a later repair exits successfully |
| Usage/observability suddenly shows no OpenCode or Codex-ledger data | The SQLite source can be absent, busy, corrupt, or query-incompatible; these are no longer collapsed into an ordinary empty result | Inspect the local-source chips at the top of the dashboard Usage area (or `sourceHealth` in usage-index JSON). A degraded OpenCode scan retains in-window last-good cached sessions; repair the named source before treating zero as observed truth |
| Observability does not show a live host process | Runtime discovery uses the numeric UID running the dashboard and is macOS/Linux-only; `sudo`, a service account, Windows, a private container PID namespace, missing `ps`/`lsof`, or restricted `/proc` changes what is visible | Run `ak dashboard` as the same ordinary OS account as the host CLI. Do not use `sudo`; use retained History on Windows and inspect OS/container process permissions when runtime presence is degraded |
| Usage suddenly shows no data for one host, or a lower total than expected | Any of the four local sources (Claude/Codex transcript roots, OpenCode's SQLite store, the Codex thread ledger) can go absent, busy, corrupt, or query-incompatible; none of these are collapsed into an ordinary empty result | Inspect the local-source chips at the top of the dashboard Usage area (or `sourceHealth` in usage-index JSON) — one chip per host; the Codex chip folds its transcript-root and thread-ledger statuses together (worse status leads, both shown in its detail text). A degraded OpenCode scan retains in-window last-good cached sessions; repair the named source before treating zero as observed truth |
| Observability does not show a live host process | Runtime discovery uses the numeric UID running the dashboard and is macOS/Linux-only; `sudo`, a service account, Windows, a private container PID namespace, missing `ps`/`lsof`, or restricted `/proc` changes what is visible | Run `ak dashboard` as the same ordinary OS account as the host CLI. Do not use `sudo`; use retained History on Windows and inspect OS/container process permissions when runtime presence is degraded. If the UID matches and none of the above applies, set `AK_RUNTIME_DEBUG=1` for one reproduction — stage-level evidence (survey row count, host classification per PID, nested-child exclusions, cwd resolution) goes to `$XDG_STATE_HOME/agentic-kit/runtime-debug.log` (mode 0600, bounded at 64 KiB; `AK_RUNTIME_DEBUG_FILE` to redirect it), then unset debug |
| Don't want the RuvNet Brain (the ~2 GB KB download) | It's on by default | `ak setup --no-ruvnet-brain`, or set `ruvnetBrain: false` in `~/.config/agentic-kit/kit.json` |
| Don't want the security surface managed | Also on by default | `ak setup --no-security` (persists `security:false`; status shows an info row and sync stops healing it) |
| RuvNet Brain KB lives somewhere non-default | The installer + ak honor `$RUVNET_BRAIN_KB` (default `~/.cache/ruvnet-brain/kb`) | export `RUVNET_BRAIN_KB` so detection points at your KB |
Expand Down
31 changes: 19 additions & 12 deletions docs/USAGE-SCORECARD-METRICS.md
Original file line number Diff line number Diff line change
Expand Up @@ -72,14 +72,21 @@ Nothing in this transcript pipeline calls a provider API or a billing endpoint;
metric is ever a copy of an actual invoice.** That is the whole reason every transcript-derived
dollar figure is labelled "API-equivalent."

The built index also exposes `sourceHealth` for the OpenCode SQLite store and
Codex thread ledger. Each source is `ok`, `absent`, `degraded`, or `not-read`,
with a bounded reason such as `busy`, `corrupt`, `query`, or `schema`. A
The built index also exposes `sourceHealth` for all four local sources: the
Claude and Codex transcript roots themselves (`claude`, `codex`), plus the two
secondary/corrective reads layered on top of them (`opencode`'s SQLite store,
`codexLedger`'s thread-attribution ledger). Each source is `ok`, `absent`,
`degraded`, or `not-read`, with a bounded reason such as an fs error code
(`ENOENT`, `EACCES`, `ENOTDIR`), `busy`, `corrupt`, `query`, or `schema`. A
degraded OpenCode read retains in-window last-good cached sessions rather than
turning an unreadable database into an observed zero. Source health is
diagnostic evidence; it is not added to token or cost totals. The dashboard
renders these states as local-source chips above every Usage view so a degraded,
absent, or deliberately unread source cannot be mistaken for healthy empty data.
renders these states as local-source chips above every Usage view, one per
HOST rather than one per field — `codex` and `codexLedger` are both Codex-only
evidence, so they fold into a single "Codex" chip carrying both sub-statuses —
so a degraded, absent, or deliberately unread source cannot be mistaken for
healthy empty data. See [ADR-0023 §7](adr/0023-fail-closed-operations-and-explicit-degradation.md)
for why the four fields are tracked to different degrees of external documentation.

The current persisted field named `provider` identifies which host transcript parser produced a
session row; it is not sufficient evidence of the inference provider. The Proposed model in
Expand Down Expand Up @@ -114,9 +121,9 @@ responses = Σ over included sessions of session.responses
**Source:**

- Filter: a parsed record with zero assistant turns is dropped entirely — "no
assistant turn → not a session" (`usage-index.mjs:904`) — and a record whose
assistant turn → not a session" (`usage-index.mjs:923`) — and a record whose
last activity falls outside the requested window is dropped too
(`usage-index.mjs:905`).
(`usage-index.mjs:924`).
- `responses` accumulation: Claude increments per assistant message
(`usage-index.mjs:504-509`); Codex increments per `agent_message` event
(`usage-index.mjs:651-655`).
Expand Down Expand Up @@ -199,7 +206,7 @@ already in effect on the given day, comparing ISO date strings
lexicographically so no `Date` parsing is involved and the module stays
clock-free.

`aggregate()` passes each usage row's own `day` (`usage-index.mjs:892`), which
`aggregate()` passes each usage row's own `day` (`usage-index.mjs:911`), which
it already has because rows are keyed by `(day, model)`. **This is the whole
point:** tokens metered in August must still read as August's rate when the
panel is opened in December. Pricing by *today's* date instead would restate a
Expand Down Expand Up @@ -281,7 +288,7 @@ tokens = input + output + cacheRead + cacheWrite (summed across all rows in wi
```

**Source:** `t.tokens` from `totals`, accumulated per row at
`usage-index.mjs:916` (`rowTokens = row.input + row.output + row.cacheRead +
`usage-index.mjs:935` (`rowTokens = row.input + row.output + row.cacheRead +
row.cacheWrite`) and rolled into `totals.tokens` via `addTo`
(`usage-index.mjs:843-852`). Rendered with `fmtTok()`
(`dashboard/client.mjs`): `≥1e9` → `"X.XB"`, `≥1e6` → `"X.XM"`,
Expand Down Expand Up @@ -453,7 +460,7 @@ session that runs from 23:58 local to 00:05 local is billed to the day its
*first* row landed on (test:
`tests/kit/usage-index.test.mjs:634`, "a session that opens before midnight
is counted on its first billed day"). Accumulation:
`byDay[row.day].cost += rowCost` (`usage-index.mjs:827`). Bar height:
`byDay[row.day].cost += rowCost` (`usage-index.mjs:846`). Bar height:
`h = maxDay ? max(2, cost/maxDay*100) : 2` (`dashboard/client.mjs`) —
every non-empty day gets a visually nonzero bar (floor of 2%), so a very
cheap day is never rendered as invisible.
Expand All @@ -474,7 +481,7 @@ renders "no sessions in window" instead of zeroed figures
(`dashboard/client.mjs`).

**Formula:** identical aggregation to every other bucket
(`byProvider[s.provider]`, populated via `addTo()`, `usage-index.mjs:870-879`,
(`byProvider[s.provider]`, populated via `addTo()`, `usage-index.mjs:889-898`,
called once per session at `usage-index.mjs:1000`), keyed by the literal string
`"claude"` or `"codex"` assigned at parse time
(`blankSession(id, 'claude')` / `blankSession(id, 'codex')`,
Expand Down Expand Up @@ -516,7 +523,7 @@ punchcard[dow + "-" + hour] += 1 per assistant/agent_message response, at its
**Source:** incremented once per Claude assistant turn
(`usage-index.mjs:504-509`, keyed by `punchKey(at)`) and once per Codex
`agent_message` (`usage-index.mjs:651-655`), merged into the window-level
`punchcard` object per session (`usage-index.mjs:981`). Cell intensity is
`punchcard` object per session (`usage-index.mjs:1000`). Cell intensity is
linear against the single busiest cell in the window:
`v = pcMax ? n/pcMax : 0` (`dashboard/client.mjs`) — this is a
**relative**, not absolute, scale, so the heatmap's brightest cell is always
Expand Down
56 changes: 51 additions & 5 deletions docs/adr/0023-fail-closed-operations-and-explicit-degradation.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,10 +2,14 @@

- **Status:** Implemented
- **Date:** 2026-08-04
- **Updated:** 2026-08-04
- **Updated:** 2026-08-06
- **Update note:** Generalized setup preflight into a required host-adapter trust contract, added
Codex registration/OpenCode approval disclosure, documented current-UID installation-mode
boundaries, and surfaced usage-source health in the dashboard UI.
boundaries, and surfaced usage-source health in the dashboard UI. Closed a parity gap §7 left
behind: `sourceHealth` covered only the two secondary/corrective sources (OpenCode's SQLite
store, Codex's thread ledger) and never the primary Claude/Codex transcript roots, so a missing
or unreadable `~/.claude/projects` or `~/.codex/sessions` still silently read as zero — the exact
failure class this ADR exists to close, just left open on the two sources most people depend on.
- **Deciders:** agentic-kit maintainers
- **Related:** [issue #111](https://github.com/pacphi/agentic-kit/issues/111),
[ADR-0008](0008-guidance-target-scope-split.md),
Expand Down Expand Up @@ -82,6 +86,15 @@ survey omits argv. A second query reads argv only for current-user executables t
host controller or Node launcher. CWD lookup then receives only filtered controller PIDs. The public
event boundary remains path-redacted as defined by ADR-0012.

`AK_RUNTIME_DEBUG=1` traces this pipeline stage-by-stage (survey row count, which PIDs were classified
as a host, which were dropped as a nested child of another candidate, and per-controller cwd
resolution) to `$XDG_STATE_HOME/agentic-kit/runtime-debug.log`, mirroring the statusline diagnostic's
opt-in/bounded-at-64-KiB/owner-only-0600 contract (`AK_RUNTIME_DEBUG_FILE` overrides the path). It is a
narrower redaction than the statusline diagnostic: raw argv/command strings are still never logged
(a pasted prompt or token could be sitting in one), but cwd paths ARE — resolving "why didn't project
X's controller show up" is the flag's entire purpose, the operator turned it on deliberately, and a
local directory path is not a secret the way a command line can be.

### 6. Every host setup has a pre-mutation trust boundary

Each host adapter must declare whether agentic-kit manages approval grants or leaves the host's
Expand All @@ -99,11 +112,41 @@ rules survive. OpenCode discloses its user-scope wildcard approvals, MCP registr
plugin, and managed host assets. Codex discloses MCP/AQE registrations while explicitly retaining
its sandbox and approval policy.

### 7. Usage source degradation is visible in the dashboard
### 7. Usage source degradation is visible in the dashboard, for all four local sources

The Usage API's `sourceHealth` field is rendered as persistent local-source chips across Usage
views. `ok`, `absent`, `degraded`, and `not-read` remain distinct, and bounded reasons such as
`busy`, `corrupt`, `query`, `schema`, or `sandboxed-roots` are visible without entering raw JSON.
`busy`, `corrupt`, `query`, `schema`, `sandboxed-roots`, or an fs error code (`ENOENT`, `EACCES`,
`ENOTDIR`) are visible without entering raw JSON.

`sourceHealth` originally covered only the two sources with a *secondary, corrective* read layered
on top of a primary parse — OpenCode's SQLite store and Codex's own thread ledger — because those
are exactly where item 2 above found silent collapse in practice. It did not cover the primary
Claude and Codex transcript roots (`~/.claude/projects`, `~/.codex/sessions`) themselves: `listClaude`
and `listCodex` walk those directories through a `readdirSync` wrapped in a bare `catch { return [] }`,
so a missing root, a permissions error, or any other I/O failure was indistinguishable from "no
sessions in this window" — the same silent-zero failure class item 2 closed for OpenCode/Codex, just
left open on the two sources every installation actually depends on.

Closing it required checking what's real to check against, not inventing a status: Anthropic
documents `~/.claude/projects/<encoded-cwd>/*.jsonl` and its 30-day default retention directly (Claude
Code's Data usage page; the `transcript_path` every hook receives). Codex's `~/.codex/sessions/**/rollout-*.jsonl`
is real and load-bearing for Codex's own `codex resume`, though OpenAI does not publish it as a
formal contract the way Anthropic does. Both are markedly more stable ground than the two sources
already tracked — the Codex thread ledger (`state_N.sqlite`) and OpenCode's `opencode.db` are
undocumented internal storage, confirmed real only by their own upstream bug trackers (e.g.
openai/codex#21750, a corrupt `state_5.sqlite` wedging Codex's own startup) and defensive code
already treating the ledger's generation suffix as "not a stable name." None of the four are
fabricated; `rootHealth()` performs a real `readdirSync` against a real path exactly like the
existing checks, just one level up the trust stack from the two sources already wired.

`sourceHealth` now reports `claude`, `codex`, `opencode`, and `codexLedger`. The dashboard renders
this by HOST, not by field: three chips for the three supported hosts (Claude, Codex, OpenCode), not
four. `codex` and `codexLedger` are both Codex-only evidence, so they fold into one "Codex" chip —
its status is the worse of the two, and both sub-statuses stay visible in the chip's detail text
(e.g. `Codex: degraded — transcripts: ok · ledger: corrupt`). No evidence is dropped; the API keeps
four independently-diagnosable fields, the UI just groups by the thing the operator actually cares
about (which host needs attention), matching how Claude and OpenCode already render as one chip each.

### 8. Clean-machine proof is isolated at every mutable boundary

Expand All @@ -124,6 +167,9 @@ setup on `macos-latest` with all global packages and user/project files under `r
undisclosed Claude project grants introduced by upstream initializers.
- Some formerly best-effort writes now fail. This is deliberate: when ak promises a backup, mutation
without one is a correctness failure.
- An unreadable `~/.claude/projects` or `~/.codex/sessions` (permissions, a corrupt filesystem entry,
the path replaced by a non-directory) now renders as a degraded local-source chip instead of a
quietly empty Usage scorecard; all four local sources share one status vocabulary.

## References

Expand All @@ -132,6 +178,6 @@ setup on `macos-latest` with all global packages and user/project files under `r
`src/lib/dashboard/{page,client,styles}.mjs`,
`src/commands/{setup,sync}.mjs`, and `src/templates/statusline-footer.cjs`.
- Tests: `tests/kit/{clean-machine-setup,heal-natives,sqlite,settings-config,blocks,
live-process-sessions,setup-command,trust-manifest,usage-index-opencode}.test.mjs`,
live-process-sessions,setup-command,trust-manifest,usage-index,usage-index-opencode}.test.mjs`,
`tests/dashboard.test.cjs`, and `tests/statusline-segments.test.cjs`.
- Clean-machine workflow: `.github/workflows/nightly.yml`.
6 changes: 5 additions & 1 deletion docs/adr/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -121,7 +121,11 @@ interop API.
managed fallbacks report degradation, SQLite retains classified failure evidence and last-good
usage, promised backups fail closed before atomic replacement, status-line failures gain redacted
opt-in diagnostics, process discovery is current-user and argv-minimized, setup discloses and
verifies its project auto-approve manifest, and clean-machine tests isolate every mutable path.
verifies its project auto-approve manifest, and clean-machine tests isolate every mutable path. A
follow-up closed a parity gap in the last item: `sourceHealth` originally covered only the two
secondary/corrective sources (OpenCode's store, Codex's thread ledger), not the primary Claude/Codex
transcript roots — an unreadable `~/.claude/projects` or `~/.codex/sessions` still read as silent
zero. It now reports all four.

**0024** gives Overview's Intelligence view real trend data instead of a permanently-empty strip:
a new `intel-history.mjs` module reads the neural pattern store, its lifetime learned-pattern
Expand Down
Loading
Loading