agentsview version
agentsview v0.41.1 (commit a902515, built 2026-08-18T13:35:15Z)
Install method
Built from source
OS / platform
macOS 26.5 (arm64)
Which agent and version
Claude Code 2.1.231, subscription auth
Which model(s)
claude models with 1hr cache
What happened, and what did you expect
Claude Code transcripts declare the cache-write TTL split per request:
"cache_creation_input_tokens": 8989,
"cache_creation": { "ephemeral_1h_input_tokens": 8989,
"ephemeral_5m_input_tokens": 0 }
However, the 5 minute rate gets used when calculating the cost above, instead of the 1 hour rate.
Here is the current sequence behind calculating:
- The Claude parser stores
message.usage verbatim, so the TTL
breakdown is already in every messages.token_usage blob.
- At query time the usage scanner reads five flat keys only and skips
the nested cache_creation object
(parseUsageTokenCountersWithReasoning, internal/db/usage.go:1310).
- The catalog keeps only
cache_creation_input_token_cost (the 5m
rate); LiteLLM's cache_creation_input_token_cost_above_1hr is
dropped at parse (internal/pricing/catalog/litellm.go:129).
- Cost = flat cache-write total x 5m rate, so 1h writes (billed 2x
input by Anthropic) are priced at 1.25x.
Sample session file or snippet
{"type":"user","cwd":"/home/user/demo","sessionId":"aaaaaaaa-bbbb-cccc-dddd-eeeeeeee0001","version":"2.1.231","uuid":"11111111-1111-1111-1111-111111111101","timestamp":"2026-08-13T12:00:00.000Z","isSidechain":false,"message":{"role":"user","content":"Reply with exactly: hello"}}
{"type":"assistant","cwd":"/home/user/demo","sessionId":"aaaaaaaa-bbbb-cccc-dddd-eeeeeeee0001","version":"2.1.231","uuid":"11111111-1111-1111-1111-111111111102","timestamp":"2026-08-13T12:00:05.000Z","isSidechain":false,"requestId":"req_demo000000000000000000001","message":{"id":"msg_demo000000000000000000001","type":"message","role":"assistant","model":"claude-fable-5","content":[{"type":"text","text":"hello"}],"stop_reason":"end_turn","usage":{"input_tokens":2,"output_tokens":62,"cache_creation_input_tokens":8989,"cache_read_input_tokens":15892,"cache_creation":{"ephemeral_1h_input_tokens":8989,"ephemeral_5m_input_tokens":0},"service_tier":"standard"}}}
{"type":"assistant","cwd":"/home/user/demo","sessionId":"aaaaaaaa-bbbb-cccc-dddd-eeeeeeee0001","version":"2.1.231","uuid":"11111111-1111-1111-1111-111111111103","timestamp":"2026-08-13T12:01:00.000Z","isSidechain":false,"requestId":"req_demo000000000000000000002","message":{"id":"msg_demo000000000000000000002","type":"message","role":"assistant","model":"claude-fable-5","content":[{"type":"text","text":"goodbye"}],"stop_reason":"end_turn","usage":{"input_tokens":2,"output_tokens":6,"cache_creation_input_tokens":77,"cache_read_input_tokens":24881,"cache_creation":{"ephemeral_1h_input_tokens":77,"ephemeral_5m_input_tokens":0},"service_tier":"standard"}}}
Steps to reproduce
Create a session, and use Claude Code's total_cost_usd as
ground truth. E.g. claude -p "Reply with exactly: hello" --output-format json.
I confirmed by checking the output of command above and viewing the cost estimated in agentsview ui.
Here's a table summarizing various ways of calculating the cost (by fable):
| Source |
Cost |
Formula |
Claude Code total_cost_usd |
$0.225533 |
cache writes at $20/MTok (1h) |
| Hand math, 1h rates |
$0.225533 |
exact |
agentsview session usage |
$0.157539 |
cache writes at $12.50/MTok (5m) |
| Hand math, 5m rates |
$0.157538 |
matches agentsview to 1 microdollar |
First request, spelled out (fable-5: input $10, output $50, cache read
$1, cache write $12.50 5m / $20 1h per MTok):
2x10 + 62x50 + 8989x20 + 15892x1 = $0.198792 (actual, matches meter)
2x10 + 62x50 + 8989x12.5 + 15892x1 = $0.131375 (agentsview)
Checklist
agentsview version
agentsview v0.41.1 (commit a902515, built 2026-08-18T13:35:15Z)
Install method
Built from source
OS / platform
macOS 26.5 (arm64)
Which agent and version
Claude Code 2.1.231, subscription auth
Which model(s)
claude models with 1hr cache
What happened, and what did you expect
Claude Code transcripts declare the cache-write TTL split per request:
However, the 5 minute rate gets used when calculating the cost above, instead of the 1 hour rate.
Here is the current sequence behind calculating:
message.usageverbatim, so the TTLbreakdown is already in every
messages.token_usageblob.the nested
cache_creationobject(
parseUsageTokenCountersWithReasoning,internal/db/usage.go:1310).cache_creation_input_token_cost(the 5mrate); LiteLLM's
cache_creation_input_token_cost_above_1hrisdropped at parse (
internal/pricing/catalog/litellm.go:129).input by Anthropic) are priced at 1.25x.
Sample session file or snippet
Steps to reproduce
Create a session, and use Claude Code's
total_cost_usdasground truth. E.g.
claude -p "Reply with exactly: hello" --output-format json.I confirmed by checking the output of command above and viewing the cost estimated in agentsview ui.
Here's a table summarizing various ways of calculating the cost (by fable):
total_cost_usdsession usageFirst request, spelled out (fable-5: input $10, output $50, cache read
$1, cache write $12.50 5m / $20 1h per MTok):
Checklist