Skip to content

Usage panel: plan limits, Codex ledger attribution, and dated pricing - #63

Merged
pacphi merged 3 commits into
mainfrom
feat/usage-limits-and-quota
Jul 27, 2026
Merged

pacphi merged 3 commits into
mainfrom
feat/usage-limits-and-quota

Conversation

@pacphi

@pacphi pacphi commented Jul 27, 2026 •

Copy link
Copy Markdown
Owner

Three related changes to the Usage panel, each with its own rationale below. All three share one theme: stop guessing at facts the data can actually answer, and stop claiming precision it cannot.

# Commit What it does Impact
1 b78d4fd Vendor-reported plan limits + Codex thread-ledger attribution New Limits sub-view; subagent double-counting fixed at the source; 18 files, +1910/−157
2 f720ac2 Excluded-turns count gets its own line Caption no longer wraps mid-phrase; layout now pinned by geometry, not markup; 4 files, +114/−37
3 e16867e Dated rate schedules, priced on the day tokens were spent Removes a CI time bomb; historical windows stop being restated when a rate changes; 7 files, +309/−54

1 · Plan limits and Codex ledger attribution (b78d4fd)

Why. The scorecard could say what tokens cost but never how much of the plan they consumed. ADR-0009 §3 excluded that deliberately: a locally computed "percentage of plan used" needs a denominator no vendor publishes, and an invented denominator is worse than no number.

That reasoning was about denominators the kit would have to fabricate. Both vendors now hand over their own percentages through supported channels, so the denominator is no longer invented and the exclusion no longer applies. ADR-0010 records the amendment.

The two admissible channels — both credential-free for ak:

Vendor Channel Mechanism
Claude statusLine push Claude Code pushes rate_limits (session/weekly/per-model used-% + reset epochs) into every statusLine invocation on Pro/Max. The managed footer tees it to claude-rate-limits.json.
Codex codex app-server One initialize → account/rateLimits/read JSON-RPC exchange with a spawned codex, which authenticates itself. Same shell-out trust model as ak status --json, TTL-cached.

The dashboard server still opens no sockets to the internet: the Claude path is a local file read, the Codex path is vendor code using vendor auth.

Explicit non-paths. No api.anthropic.com/api/oauth/usage — undocumented, hostile to unrecognized clients, and consumer-OAuth use outside Claude Code is ToS-prohibited and server-enforced since January 2026. No Keychain reads (the macOS item stopped carrying the token in Claude Code 2.1.x). No chatgpt.com/backend-api/wham/*. No auto-consumption of Codex reset credits — the panel reports them; redeeming a finite grant is the user's call.

The trap that shaped the data model. Codex's primary/secondary window fields do not reliably mean 5-hour/weekly. A live prolite account answered with primary.windowDurationMins = 10080 — the weekly — and secondary: null. Every window is keyed and labelled by duration, never by slot. Trusting the field name would have mislabelled every bar.

Codex attribution stops being heuristic. Codex keeps its own thread ledger (state_N.sqlite — the suffix is a migration generation, so it is globbed) carrying thread_source and thread_spawn_edges. A ledger-identified subagent now has its usage stripped, because its rollout replays the parent's entire token history — ccusage/ccusage#950 measured up to 91× inflation. The rollout sniff from #60 remains the fallback. Rollouts also yield reasoning tokens (a subset of output, annotation only) and embedded rate-limit snapshots, making a utilization history reconstructable with zero network calls.

Findings. Six new detectors under the existing evidence discipline — vendor percentages count as the user's own data, no dollar impact is ever claimed from a percentage, and "now" is the payload's generatedAt rather than a clock, so every firing is reproducible: limit-pacing, cross-host-arbitrage, codex-reset-credits, plus local parity with Claude's own /usage characteristics (parallel-sessions, subagent-share, long-session-share).

Also fixed: classify-coverage advertised ak x usage classify --enrich, which does not exist. A dead command in a diagnostic is worse than none.

UI clarity. The Providers strip is retitled routed models (it collided with Usage's "models in play") and explains itself: it is the per-activity policy projected into agentic-qe agent overrides and ak dual run — not a record of what ran. agentic-qe carries its own model router, so a route there is an assignment, not a guarantee. Internal ADR ids were removed from all visible text, and the UI net now fails on any that reappear.

2 · The excluded-turns count gets its own line (f720ac2)

As an inline tail on the models caption it wrapped mid-phrase, leaving · 4 dangling at the end of one line and "dropped/errored turns excluded" orphaned on the next — the count read as part of the caption rather than as its own fact. It is now a block sub-line with no leading separator (a · starting a line reads as a continuation).

The harness could not have caught this: the fixture corpus contained no errored turn, so the element rendered empty and the wrap never happened. It now carries an isApiErrorMessage turn, and six assertions pin the layout by geometry — comparing the sub-line's rect against a Range measured over the caption text preceding it — because markup alone cannot prove a reader sees two distinct lines.

3 · Dated rate schedules (e16867e)

Why. Sonnet 5's introductory rate reverts to $3/$15 on 2026-09-01. That was a code comment promising a human would edit it, enforced by a test that read the wall clock and failed CI on a calendar date. A build that goes red for a reason that is not a defect teaches people to ignore builds.

Clock-switching would have been the wrong fix. Cost attribution is historical — tokens metered in August must still read as August's rate in December. Selecting by "now" restates finished windows the moment a published rate changes: a Sonnet-heavy August would jump 50% overnight with no session having changed, under a panel claiming to show "what these tokens would cost metered".

So the rate is selected by the day the tokens were spent, which aggregate() already had in hand and was simply not passing:

Day Sonnet 5 rate 1M in + 1M out
2026-08-31 $2/$10 $12.00
2026-09-01 $3/$15 $18.00

Uniform across providers, because a date range is a fact about a price, not about a vendor. anthropic(5, 25) and openai(2.5, 15) build one-period schedules and read exactly as before; .dated([...]) exists identically on both. OpenAI publishes no dated promos today, so its entries are all single-period — a fact about the data, not the table's shape. A Codex promo is a one-line edit rather than a second mechanism built under deadline with its own boundary tests to drift. A test asserts no provider-specific rate shape exists.

Deliberately not absorbed: rates varying by how a request was served — regional uplift, large-prompt surcharge, service tiers. Transcripts record neither endpoint nor tier; a general modifier system would manufacture precision the data cannot support. They stay in UNMODELLED_PRICING_FACTORS.

A subtlety a test caught. A dateless priceFor prices as of PRICES_AS_OF, not the newest period. "Newest" only means "current" once every published change has landed, and judging that needs a clock this module does not read. The verification date is also what the UI prints as "rates as of …", so the default and the label cannot disagree. My first version defaulted to the newest period and an existing test expecting today's $2/$10 failed — the test doing its job.


Verification

  • pnpm run check exits 0 — 569 unit tests + all cjs suites, including 57 new tests and the self-verifying file:line doc citations (re-anchored twice as these changes shifted them; the net caught both, which is its entire purpose).
  • Playwright artifact net 108/108, extended to the Limits view with a deliberately null-heavy stub — it caught a real bug where a null reset time rendered as an epoch-1970 date.
  • /api/limits smoke-tested against real local state: prolite plan, both lanes, 2 reset credits, a genuine codex-reset-credits finding.

Docs. ADR-0010 is new; ADR-0009 §3 records both the quota amendment and the day-of-spend pricing decision; USAGE-SCORECARD-METRICS gains §3a, §13b, §13c. PROVIDERS.md is deliberately untouched — it covers host/provider configuration, not rate mechanics or quota reads.

Note for reviewers: Claude's limit bars read "no data" until ak sync injects the updated statusline footer and one Claude session runs — the tee is push-only by design, and the empty state says so. Codex's bars populate immediately.

pacphi added 3 commits July 27, 2026 08:21
The scorecard could say what tokens cost but never how much of the plan
they consumed — ADR-0009 §3 excluded that deliberately, because a locally
computed percentage needs a denominator no vendor publishes. Both vendors
now hand over their own percentages through supported channels, so the
denominator is no longer invented and the exclusion no longer applies.

ADR-0010 defines the two admissible channels, both credential-free for ak:

  claude — Claude Code PUSHES rate_limits (session/weekly/per-model used
    percentage + reset epochs) into every statusLine invocation on Pro/Max.
    The managed footer tees that payload to claude-rate-limits.json.
  codex — one initialize -> account/rateLimits/read exchange with a spawned
    `codex app-server`, which authenticates itself. Same shell-out trust
    model as `ak status --json`; TTL-cached.

Explicit non-paths: no /api/oauth/usage (undocumented, and consumer-OAuth
use outside Claude Code is ToS-prohibited and server-enforced since Jan
2026), no Keychain reads, no chatgpt.com backend endpoints, and no
auto-consumption of Codex reset credits.

Windows are keyed by DURATION, never by the vendor's primary/secondary
slot: a live prolite account reported `primary` as the 10080-minute
weekly window. Field-name trust would have mislabelled every bar.

Codex attribution stops being heuristic. Codex keeps its own thread ledger
(state_N.sqlite — the suffix is a migration generation, so glob it) with
thread_source and spawn edges; a ledger-identified subagent has its usage
stripped, since its rollout replays the parent's whole token history
(ccusage#950 measured up to 91x inflation). The rollout sniff from #60
remains the fallback. Rollouts also yield reasoning tokens (a subset of
output — annotation only, never summed) and embedded rate-limit snapshots,
making a utilization history reconstructable with zero network.

Six new detectors under the existing evidence rules — vendor percentages
are the user's own data, and no dollar impact is ever claimed from a
percentage. "Now" is the payload's generatedAt, never a clock, so every
firing is reproducible from its input.

Also: the classify-coverage finding advertised `ak x usage classify
--enrich`, which does not exist — a dead command in a diagnostic is worse
than none. And the Sonnet 5 introductory price was a comment promising a
2026-09-01 revert with nothing enforcing it; it is now a test that fails
the suite from that date until the table is corrected.

The Providers strip is retitled "routed models" and explains itself: it is
the per-activity policy projected into agentic-qe agent overrides and
`ak dual run`, not a record of what ran. agentic-qe carries its own model
router, so a route there is an assignment, not a guarantee.

Verified: pnpm run check green (569 unit + all cjs suites, incl. 47 new
tests and the self-verifying doc citations, re-anchored after drift);
Playwright artifact net 102/102, now also failing on any visible ADR id.
As an inline tail on the models caption it wrapped mid-phrase, leaving
"· 4" dangling at the end of one line and "dropped/errored turns
excluded" orphaned on the next — the count read as part of the caption
rather than as its own fact. It is now a block sub-line with no leading
separator (a "·" at the start of a line reads as a continuation).

The UI harness could not have caught this: the fixture corpus contained
no errored turn, so the element rendered empty and the wrap never
happened. It now carries an isApiErrorMessage turn, and six assertions
pin the layout by GEOMETRY — comparing the sub-line's rect against a
Range measured over the caption text that precedes it — because markup
alone cannot prove a reader sees two distinct lines.

Doc citations re-anchored after the CSS insertion shifted them.
Every table entry is now a schedule — an ordered list of periods, each
with the day it takes effect — and priceFor(model, provider, day) picks
the period in effect on that day. The single-rate case is a one-period
schedule, so `anthropic(5, 25)` reads exactly as before.

This replaces a comment promising a human would edit Sonnet 5 on
2026-09-01, and the wall-clock test that enforced the promise by failing
CI on a calendar date. A build that goes red for a reason that is not a
defect teaches people to ignore builds.

Clock-switching would have been the wrong fix. Cost attribution is
historical: tokens metered in August must still read as August's rate in
December. Selecting by "now" restates finished windows the moment a
published rate changes — a Sonnet-heavy August would jump 50% overnight
with no session having changed, under a panel that claims to show "what
these tokens would cost metered". So the day comes from the usage row,
which aggregate() already had in hand and was simply not passing.

The mechanism is identical for both providers, because a date range is a
fact about a price, not about a vendor. OpenAI publishes no dated promos
today, so every OpenAI entry is one period — but that is a fact about the
DATA, not a gap in the table. openai.dated([...]) exists and behaves the
same, so a Codex promo is a one-line edit rather than new machinery built
under deadline pressure with a second set of boundary tests to keep in
sync. This mirrors costOf, which has no per-provider branch for the same
reason.

Deliberately NOT absorbed: rates that vary by how a request was served —
regional uplift, large-prompt surcharge, service tiers. Those are a
different axis, transcripts record neither endpoint nor tier, and
stretching schedules to cover them would manufacture precision the data
cannot support. They stay in UNMODELLED_PRICING_FACTORS.

One subtlety worth stating: a dateless priceFor prices as of
PRICES_AS_OF, not the newest period. "Newest" only means "current" once
every published change has landed, and judging that needs a clock this
module does not read. The verification date is also what the UI already
prints as "rates as of ...", so the default and the label cannot
disagree. Caught by an existing test that expected today's $2/$10.

Tests: the time bomb is now ten deterministic boundary assertions,
including the property that matters — a finished window is not restated
when a later change takes effect. A seam test pins that aggregate()
actually forwards row.day, since the schedule is inert if it does not.

Docs: ADR-0009 §3 records the day-of-spend decision and its two bounding
rules; USAGE-SCORECARD-METRICS gains §3a and ties the two Sonnet rows in
the §13 rate table to the two periods the code encodes. PROVIDERS.md is
untouched — it covers host/provider configuration, not rate mechanics.
@pacphi pacphi changed the title Usage panel: vendor-reported plan limits, Codex ledger attribution Usage panel: plan limits, Codex ledger attribution, and dated pricing Jul 27, 2026
@pacphi
pacphi merged commit 13ae952 into main Jul 27, 2026
11 checks passed
@pacphi
pacphi deleted the feat/usage-limits-and-quota branch July 27, 2026 15:45
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant