Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
106 changes: 53 additions & 53 deletions docs/TRANSCRIPTS.md

Large diffs are not rendered by default.

264 changes: 187 additions & 77 deletions docs/USAGE-SCORECARD-METRICS.md

Large diffs are not rendered by default.

29 changes: 29 additions & 0 deletions docs/adr/0009-usage-scorecard-local-transcript-analytics.md
Original file line number Diff line number Diff line change
Expand Up @@ -65,11 +65,40 @@ single-flight: a refresh already in progress is joined, not duplicated.
`src/lib/pricing.mjs` holds dated per-model rates and computes
`input×rate + cacheWrite×rate×1.25 + cacheRead×rate×0.1 + output×outRate`.

**Rates are a schedule, and a row is priced on the day it was spent.** Every table entry is an
ordered list of periods (the common case being one period that has always applied), and
`priceFor(model, provider, day)` selects the period in effect on that day. Cost attribution is
historical: tokens metered in August must still read as August's rate in December. Switching rates
by *today's* date instead would retroactively restate finished windows the moment a published rate
changed — which a panel claiming "what these tokens would cost metered" cannot do. Two constraints
follow, both enforced in code:

- **Only published changes may be encoded.** A schedule records a rate change the vendor has
announced (Anthropic's Sonnet 5 introductory period ending 2026-08-31). Encoding a *forecast*
would fabricate data, the same error as the invented denominator above.
- **The mechanism is identical for both providers**, because a date range is a fact about a price,
not about a vendor. That OpenAI currently publishes no dated promos is a fact about the data, not
a reason for a second code path — mirroring how `costOf` needs no per-provider branch. Rates that
vary by *how* a request was served (regional uplift, large-prompt surcharge, service tiers) are a
different axis, are not expressible as a schedule, and stay in `UNMODELLED_PRICING_FACTORS`
because transcripts do not record the endpoint or tier.

A dateless `priceFor` prices as of `PRICES_AS_OF`, the table's verification date — not the newest
period, which only becomes "current" once every published change has landed, a judgement requiring
a clock this module deliberately does not read.

On a Max/Pro subscription the user is **not billed per token**, and Codex-via-ChatGPT likewise. The
panel therefore labels every figure **"API-equivalent — not your plan billing"** as a first-class UI
element, not a footnote. We do not model plan utilisation: Anthropic does not publish the limits that
would make such a percentage honest, and an invented denominator is worse than no number.

> **Amended by [ADR-0010](0010-provider-mediated-quota-reads.md) (2026-07-27):** the exclusion above
> was about denominators the kit would have to *invent*. Both vendors now hand over their own
> percentages through supported channels (Claude Code's statusLine `rate_limits` push; Codex's
> `app-server` RPC), and the Limits sub-view renders those vendor-reported figures — with
> provenance and freshness — under ADR-0010's provider-mediated rules. Locally *computed*
> plan percentages remain excluded, exactly as stated here.

OpenAI publishes no pricing in `~/.codex/models_cache.json` (verified), so Codex rates are a
maintained table and are **date-stamped** in the UI so staleness is visible rather than silent.

Expand Down
98 changes: 98 additions & 0 deletions docs/adr/0010-provider-mediated-quota-reads.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,98 @@
# ADR-0010 — Provider-mediated quota reads (the Limits view)

Date: 2026-07-27 · Status: **Accepted** · Amends: ADR-0009 §3

## Context

ADR-0009 §3 deliberately excluded plan/limit modeling from the usage scorecard:
Anthropic and OpenAI publish no limit values, so any locally computed
"percentage of plan used" would rest on an invented denominator — and "an
invented denominator is worse than no number". That reasoning was about
denominators the kit would have to *fabricate*. Research (2026-07-27, cited in
§Research below) established that both vendors now *hand over* their own
percentages through supported channels:

- **Claude Code** pushes a `rate_limits` object — `five_hour` / `seven_day`
(and per-model weekly buckets) with `used_percentage` and `resets_at` — into
every statusLine invocation for Pro/Max subscribers. This is documented
behavior, not an internal.
- **Codex** answers `account/rateLimits/read` over its `codex app-server`
JSON-RPC surface with per-lane (`rateLimitsByLimitId`) used-percent, window
duration, reset time, plan type, and rate-limit reset credits.

A vendor-reported percentage IS an honest denominator. What remains dishonest —
and stays excluded — is fetching those numbers through channels the vendors
prohibit or do not support.

## Decision

Quota data enters the kit **only through the vendors' own software**, with ak
never reading, storing, or refreshing a vendor credential:

1. **Claude — statusline tee.** The kit's managed statusline footer tees the
pushed `rate_limits` (plus `context_window` and `cost`) to
`~/.config/agentic-kit/claude-rate-limits.json` (0600, atomic, throttled to
one write per minute). Push, not pull: with no recent Claude session the
file goes stale, and the UI labels it stale rather than hiding it.
2. **Codex — app-server subprocess.** `src/lib/quota.mjs` spawns
`codex -s read-only -a untrusted app-server` (codex authenticates itself),
performs one `initialize` → `account/rateLimits/read` exchange with a hard
timeout and kill, and caches the normalized answer for `CODEX_TTL_MS`.
This is the same shell-out trust model the dashboard already uses for
`ak status --json`.
3. **The dashboard server itself still opens no sockets to the internet.**
ADR-0005/0007's egress split is unchanged: the Codex subprocess is vendor
code using vendor auth, and the Claude path is a local file read.

Normalization rules (all in `quota.mjs`, all pinned by tests):

- **Windows are keyed by duration, never by slot.** Codex's `primary` /
`secondary` fields do not reliably mean "5-hour" / "weekly" — a live
`prolite` account answered with `primary.windowDurationMins = 10080`.
- `null` timestamps stay `null`; they are never coerced to epoch 0.
- Every payload carries `fetchedAt`, and the UI renders freshness ("as of Nm
ago") next to every number.

### Explicit non-paths

- **No `api.anthropic.com/api/oauth/usage`.** Undocumented; hostile
rate-limiting to unrecognized clients; and Anthropic's consumer ToS bars
subscription OAuth tokens "in any other product, tool, or service", enforced
server-side since January 2026.
- **No Keychain or credential-file reads** (the macOS item stopped carrying
the OAuth token in Claude Code 2.1.x; refresh tokens rotate destructively).
- **No `chatgpt.com/backend-api/wham/*`** — private endpoints that would
require ak to hold and refresh the user's bearer token.
- **No auto-consumption of Codex reset credits.** The panel reports them;
redeeming a finite grant is the user's action in codex's own `/usage`.

## Consequences

- The Limits sub-view can show authoritative, cross-device session/weekly
utilization — data local transcript parsing can never produce (ADR-0009 §3's
table of unknowables shrinks to: extra-usage credit balance and subscription
tier, which have no supported channel).
- `usage-insights.mjs` gains limit-aware detectors (`detectLimitInsights`)
under the same evidence rules; vendor percentages count as the user's own
data, and no dollar impact is ever claimed from a percentage.
- Codex attribution stops being heuristic: `codex-state.mjs` reads Codex's own
SQLite thread ledger (`state_*.sqlite`, glob — the numeric suffix is a
migration generation) for `thread_source` and spawn edges, demoting the
rollout-sniffing subagent guard to a fallback.
- Two new 0600 cache files exist under `~/.config/agentic-kit/`; both are
derived, deletable, and self-healing.
- The `codex app-server` surface is flagged experimental upstream; the client
pins nothing beyond two method names and degrades to "no data" with the
stale cache visible if the protocol shifts.

## Research

Grounded findings behind this ADR (full citations in the planning record):
statusLine `rate_limits` — code.claude.com/docs/en/statusline.md; Admin
Usage/Cost + Claude Code Analytics APIs are org-only —
platform.claude.com/docs/en/manage-claude/usage-cost-api; consumer-OAuth
prohibition — code.claude.com/docs/en/legal-and-compliance; Codex app-server
schema — `codex app-server generate-json-schema` (verified live on codex
0.145.0, 2026-07-27); Codex reset-credit detail — openai/codex#29618 (shipped
via PR #30488); subagent rollout replay inflation — ccusage/ccusage#950;
wham endpoint traffic complaints — openai/codex#10869, #27952.
79 changes: 79 additions & 0 deletions src/lib/codex-state.mjs
Original file line number Diff line number Diff line change
@@ -0,0 +1,79 @@
// Codex's own thread ledger (~/.codex/state_N.sqlite) — the authoritative
// source for per-thread attribution that the rollout JSONL can only guess at.
//
// Codex ≥0.140 maintains a `threads` table with a pre-aggregated `tokens_used`
// column plus `thread_source` ('user' | 'subagent') and `source`
// ('cli' | 'exec' | 'subagent'), and a `thread_spawn_edges` table recording
// parent→child delegation. That replaces the kit's heuristic subagent-replay
// guard (a `session_meta.thread_source` sniff, PR #60) with Codex's own
// bookkeeping: a subagent rollout replays its parent's ENTIRE token history
// (ccusage/ccusage#950 measured up to 91x inflation), so knowing which threads
// are subagents is what keeps the totals honest.
//
// Deliberate limits:
// - READ-ONLY, always. `withDb` opens readonly and swallows every error —
// a locked or half-migrated db yields null, never a crash and never a lock
// held against the codex CLI itself.
// - The filename suffix (`state_5`) is a MIGRATION GENERATION, not a stable
// name. We glob `state_*.sqlite` and take the highest generation rather
// than hardcoding today's.
// - Column names are probed before use: a future migration that drops or
// renames a column degrades this module to null (callers fall back to the
// JSONL heuristic), not to a throw.
import fs from 'node:fs';
import path from 'node:path';
import { codexDir } from './paths.mjs';
import { withDb } from './sqlite.mjs';

/** Newest-generation state db file under ~/.codex, or null. Exported for test
* via the `dir` override. */
export function codexStateDb(dir = codexDir()) {
let names;
try { names = fs.readdirSync(dir); } catch { return null; }
const gens = names
.map((n) => /^state_(\d+)\.sqlite$/.exec(n))
.filter(Boolean)
.map((m) => ({ name: m[0], gen: Number(m[1]) }))
.sort((a, b) => b.gen - a.gen);
return gens.length ? path.join(dir, gens[0].name) : null;
}

/**
* Read Codex's thread ledger. Returns
* { threads: Map<id, {tokensUsed, threadSource, source, model, gitBranch}>,
* parents: Map<childId, parentId> }
* or null when the db is absent, unreadable, or shaped unexpectedly.
*
* @param {{ dir?: string, file?: string }} [opts] test seams
*/
export function readCodexState(opts = {}) {
const file = opts.file ?? codexStateDb(opts.dir);
if (!file) return null;
return withDb(file, (db) => {
const cols = new Set(db.prepare('PRAGMA table_info(threads)').all().map((c) => c.name));
// id + thread_source are what the attribution fix rests on; without them
// this ledger cannot answer the question and the caller must fall back.
if (!cols.has('id') || !cols.has('thread_source')) return null;
const pick = ['id', 'thread_source']
.concat(['tokens_used', 'source', 'model', 'git_branch'].filter((c) => cols.has(c)));
const threads = new Map();
for (const row of db.prepare(`SELECT ${pick.join(', ')} FROM threads`).all()) {
threads.set(String(row.id), {
tokensUsed: Number(row.tokens_used) || 0,
threadSource: typeof row.thread_source === 'string' ? row.thread_source : null,
source: typeof row.source === 'string' ? row.source : null,
model: typeof row.model === 'string' ? row.model : null,
gitBranch: typeof row.git_branch === 'string' ? row.git_branch : null,
});
}
const parents = new Map();
try {
for (const e of db.prepare('SELECT parent_thread_id, child_thread_id FROM thread_spawn_edges').all()) {
if (e.child_thread_id != null && e.parent_thread_id != null) {
parents.set(String(e.child_thread_id), String(e.parent_thread_id));
}
}
} catch { /* the edges table is additive detail — attribution works without it */ }
return { threads, parents };
}, null);
}
Loading
Loading