Usage panel: plan limits, Codex ledger attribution, and dated pricing - #63
Merged
Merged
Conversation
The scorecard could say what tokens cost but never how much of the plan
they consumed — ADR-0009 §3 excluded that deliberately, because a locally
computed percentage needs a denominator no vendor publishes. Both vendors
now hand over their own percentages through supported channels, so the
denominator is no longer invented and the exclusion no longer applies.
ADR-0010 defines the two admissible channels, both credential-free for ak:
claude — Claude Code PUSHES rate_limits (session/weekly/per-model used
percentage + reset epochs) into every statusLine invocation on Pro/Max.
The managed footer tees that payload to claude-rate-limits.json.
codex — one initialize -> account/rateLimits/read exchange with a spawned
`codex app-server`, which authenticates itself. Same shell-out trust
model as `ak status --json`; TTL-cached.
Explicit non-paths: no /api/oauth/usage (undocumented, and consumer-OAuth
use outside Claude Code is ToS-prohibited and server-enforced since Jan
2026), no Keychain reads, no chatgpt.com backend endpoints, and no
auto-consumption of Codex reset credits.
Windows are keyed by DURATION, never by the vendor's primary/secondary
slot: a live prolite account reported `primary` as the 10080-minute
weekly window. Field-name trust would have mislabelled every bar.
Codex attribution stops being heuristic. Codex keeps its own thread ledger
(state_N.sqlite — the suffix is a migration generation, so glob it) with
thread_source and spawn edges; a ledger-identified subagent has its usage
stripped, since its rollout replays the parent's whole token history
(ccusage#950 measured up to 91x inflation). The rollout sniff from #60
remains the fallback. Rollouts also yield reasoning tokens (a subset of
output — annotation only, never summed) and embedded rate-limit snapshots,
making a utilization history reconstructable with zero network.
Six new detectors under the existing evidence rules — vendor percentages
are the user's own data, and no dollar impact is ever claimed from a
percentage. "Now" is the payload's generatedAt, never a clock, so every
firing is reproducible from its input.
Also: the classify-coverage finding advertised `ak x usage classify
--enrich`, which does not exist — a dead command in a diagnostic is worse
than none. And the Sonnet 5 introductory price was a comment promising a
2026-09-01 revert with nothing enforcing it; it is now a test that fails
the suite from that date until the table is corrected.
The Providers strip is retitled "routed models" and explains itself: it is
the per-activity policy projected into agentic-qe agent overrides and
`ak dual run`, not a record of what ran. agentic-qe carries its own model
router, so a route there is an assignment, not a guarantee.
Verified: pnpm run check green (569 unit + all cjs suites, incl. 47 new
tests and the self-verifying doc citations, re-anchored after drift);
Playwright artifact net 102/102, now also failing on any visible ADR id.
As an inline tail on the models caption it wrapped mid-phrase, leaving "· 4" dangling at the end of one line and "dropped/errored turns excluded" orphaned on the next — the count read as part of the caption rather than as its own fact. It is now a block sub-line with no leading separator (a "·" at the start of a line reads as a continuation). The UI harness could not have caught this: the fixture corpus contained no errored turn, so the element rendered empty and the wrap never happened. It now carries an isApiErrorMessage turn, and six assertions pin the layout by GEOMETRY — comparing the sub-line's rect against a Range measured over the caption text that precedes it — because markup alone cannot prove a reader sees two distinct lines. Doc citations re-anchored after the CSS insertion shifted them.
Every table entry is now a schedule — an ordered list of periods, each with the day it takes effect — and priceFor(model, provider, day) picks the period in effect on that day. The single-rate case is a one-period schedule, so `anthropic(5, 25)` reads exactly as before. This replaces a comment promising a human would edit Sonnet 5 on 2026-09-01, and the wall-clock test that enforced the promise by failing CI on a calendar date. A build that goes red for a reason that is not a defect teaches people to ignore builds. Clock-switching would have been the wrong fix. Cost attribution is historical: tokens metered in August must still read as August's rate in December. Selecting by "now" restates finished windows the moment a published rate changes — a Sonnet-heavy August would jump 50% overnight with no session having changed, under a panel that claims to show "what these tokens would cost metered". So the day comes from the usage row, which aggregate() already had in hand and was simply not passing. The mechanism is identical for both providers, because a date range is a fact about a price, not about a vendor. OpenAI publishes no dated promos today, so every OpenAI entry is one period — but that is a fact about the DATA, not a gap in the table. openai.dated([...]) exists and behaves the same, so a Codex promo is a one-line edit rather than new machinery built under deadline pressure with a second set of boundary tests to keep in sync. This mirrors costOf, which has no per-provider branch for the same reason. Deliberately NOT absorbed: rates that vary by how a request was served — regional uplift, large-prompt surcharge, service tiers. Those are a different axis, transcripts record neither endpoint nor tier, and stretching schedules to cover them would manufacture precision the data cannot support. They stay in UNMODELLED_PRICING_FACTORS. One subtlety worth stating: a dateless priceFor prices as of PRICES_AS_OF, not the newest period. "Newest" only means "current" once every published change has landed, and judging that needs a clock this module does not read. The verification date is also what the UI already prints as "rates as of ...", so the default and the label cannot disagree. Caught by an existing test that expected today's $2/$10. Tests: the time bomb is now ten deterministic boundary assertions, including the property that matters — a finished window is not restated when a later change takes effect. A seam test pins that aggregate() actually forwards row.day, since the schedule is inert if it does not. Docs: ADR-0009 §3 records the day-of-spend decision and its two bounding rules; USAGE-SCORECARD-METRICS gains §3a and ties the two Sonnet rows in the §13 rate table to the two periods the code encodes. PROVIDERS.md is untouched — it covers host/provider configuration, not rate mechanics.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Three related changes to the Usage panel, each with its own rationale below. All three share one theme: stop guessing at facts the data can actually answer, and stop claiming precision it cannot.
b78d4fdf720ac2e16867e1 · Plan limits and Codex ledger attribution (
b78d4fd)Why. The scorecard could say what tokens cost but never how much of the plan they consumed. ADR-0009 §3 excluded that deliberately: a locally computed "percentage of plan used" needs a denominator no vendor publishes, and an invented denominator is worse than no number.
That reasoning was about denominators the kit would have to fabricate. Both vendors now hand over their own percentages through supported channels, so the denominator is no longer invented and the exclusion no longer applies. ADR-0010 records the amendment.
The two admissible channels — both credential-free for ak:
rate_limits(session/weekly/per-model used-% + reset epochs) into every statusLine invocation on Pro/Max. The managed footer tees it toclaude-rate-limits.json.codex app-serverinitialize→account/rateLimits/readJSON-RPC exchange with a spawned codex, which authenticates itself. Same shell-out trust model asak status --json, TTL-cached.The dashboard server still opens no sockets to the internet: the Claude path is a local file read, the Codex path is vendor code using vendor auth.
Explicit non-paths. No
api.anthropic.com/api/oauth/usage— undocumented, hostile to unrecognized clients, and consumer-OAuth use outside Claude Code is ToS-prohibited and server-enforced since January 2026. No Keychain reads (the macOS item stopped carrying the token in Claude Code 2.1.x). Nochatgpt.com/backend-api/wham/*. No auto-consumption of Codex reset credits — the panel reports them; redeeming a finite grant is the user's call.The trap that shaped the data model. Codex's
primary/secondarywindow fields do not reliably mean 5-hour/weekly. A liveproliteaccount answered withprimary.windowDurationMins = 10080— the weekly — andsecondary: null. Every window is keyed and labelled by duration, never by slot. Trusting the field name would have mislabelled every bar.Codex attribution stops being heuristic. Codex keeps its own thread ledger (
state_N.sqlite— the suffix is a migration generation, so it is globbed) carryingthread_sourceandthread_spawn_edges. A ledger-identified subagent now has its usage stripped, because its rollout replays the parent's entire token history — ccusage/ccusage#950 measured up to 91× inflation. The rollout sniff from #60 remains the fallback. Rollouts also yield reasoning tokens (a subset of output, annotation only) and embedded rate-limit snapshots, making a utilization history reconstructable with zero network calls.Findings. Six new detectors under the existing evidence discipline — vendor percentages count as the user's own data, no dollar impact is ever claimed from a percentage, and "now" is the payload's
generatedAtrather than a clock, so every firing is reproducible:limit-pacing,cross-host-arbitrage,codex-reset-credits, plus local parity with Claude's own/usagecharacteristics (parallel-sessions,subagent-share,long-session-share).Also fixed:
classify-coverageadvertisedak x usage classify --enrich, which does not exist. A dead command in a diagnostic is worse than none.UI clarity. The Providers strip is retitled routed models (it collided with Usage's "models in play") and explains itself: it is the per-activity policy projected into agentic-qe agent overrides and
ak dual run— not a record of what ran. agentic-qe carries its own model router, so a route there is an assignment, not a guarantee. Internal ADR ids were removed from all visible text, and the UI net now fails on any that reappear.2 · The excluded-turns count gets its own line (
f720ac2)As an inline tail on the models caption it wrapped mid-phrase, leaving
· 4dangling at the end of one line and "dropped/errored turns excluded" orphaned on the next — the count read as part of the caption rather than as its own fact. It is now a block sub-line with no leading separator (a·starting a line reads as a continuation).The harness could not have caught this: the fixture corpus contained no errored turn, so the element rendered empty and the wrap never happened. It now carries an
isApiErrorMessageturn, and six assertions pin the layout by geometry — comparing the sub-line's rect against aRangemeasured over the caption text preceding it — because markup alone cannot prove a reader sees two distinct lines.3 · Dated rate schedules (
e16867e)Why. Sonnet 5's introductory rate reverts to $3/$15 on 2026-09-01. That was a code comment promising a human would edit it, enforced by a test that read the wall clock and failed CI on a calendar date. A build that goes red for a reason that is not a defect teaches people to ignore builds.
Clock-switching would have been the wrong fix. Cost attribution is historical — tokens metered in August must still read as August's rate in December. Selecting by "now" restates finished windows the moment a published rate changes: a Sonnet-heavy August would jump 50% overnight with no session having changed, under a panel claiming to show "what these tokens would cost metered".
So the rate is selected by the day the tokens were spent, which
aggregate()already had in hand and was simply not passing:Uniform across providers, because a date range is a fact about a price, not about a vendor.
anthropic(5, 25)andopenai(2.5, 15)build one-period schedules and read exactly as before;.dated([...])exists identically on both. OpenAI publishes no dated promos today, so its entries are all single-period — a fact about the data, not the table's shape. A Codex promo is a one-line edit rather than a second mechanism built under deadline with its own boundary tests to drift. A test asserts no provider-specific rate shape exists.Deliberately not absorbed: rates varying by how a request was served — regional uplift, large-prompt surcharge, service tiers. Transcripts record neither endpoint nor tier; a general modifier system would manufacture precision the data cannot support. They stay in
UNMODELLED_PRICING_FACTORS.A subtlety a test caught. A dateless
priceForprices as ofPRICES_AS_OF, not the newest period. "Newest" only means "current" once every published change has landed, and judging that needs a clock this module does not read. The verification date is also what the UI prints as "rates as of …", so the default and the label cannot disagree. My first version defaulted to the newest period and an existing test expecting today's $2/$10 failed — the test doing its job.Verification
pnpm run checkexits 0 — 569 unit tests + all cjs suites, including 57 new tests and the self-verifyingfile:linedoc citations (re-anchored twice as these changes shifted them; the net caught both, which is its entire purpose)./api/limitssmoke-tested against real local state:proliteplan, both lanes, 2 reset credits, a genuinecodex-reset-creditsfinding.Docs. ADR-0010 is new; ADR-0009 §3 records both the quota amendment and the day-of-spend pricing decision; USAGE-SCORECARD-METRICS gains §3a, §13b, §13c.
PROVIDERS.mdis deliberately untouched — it covers host/provider configuration, not rate mechanics or quota reads.Note for reviewers: Claude's limit bars read "no data" until
ak syncinjects the updated statusline footer and one Claude session runs — the tee is push-only by design, and the empty state says so. Codex's bars populate immediately.