Skip to content

feat(ci): ship the one benchmark input that was not committed, and guard the rest - #64

Merged
DanielWLiu07 merged 1 commit into
mainfrom
feat/ship-the-benchmark-inputs
Sep 2, 2026
Merged

DanielWLiu07 merged 1 commit into
mainfrom
feat/ship-the-benchmark-inputs

Conversation

@DanielWLiu07

Copy link
Copy Markdown
Owner

docs/bench/allocator.md's live row — parse 2.0 allocs/msg at 574 B, books 0.9 — was measured on captures/live-poly.feedlog.

/captures/ is gitignored as "large, regenerated". The file has never been in the repo.

So that row worked perfectly for anyone with this working tree and could not be reproduced by anyone else. Every command in the doc runs. The numbers come out exactly as printed. None of it is checkable from a clone.

It's the only figure in docs/bench/ in that state, against a README rule that a figure is either measured on a committed capture or labelled as not yet measured. It was neither.

Fix

The capture is 85 KB gzipped and now sits beside the other six, with the gunzip-and-replay command in the doc. Reproduces the row exactly:

records   877 (kalshi 0, polymarket 877)
allocs    parse 1796 (2.0/msg, 574 B/msg), books 784 (0.9/msg)

Guard

scripts/check_bench_inputs.py in CI: fails if any doc under docs/bench/ names a .feedlog that git doesn't track.

Angle-bracket metavariables are excluded — basis xvenue-lead <capture.feedlog> names an argument, not a file. The first version of the check flagged exactly that and was wrong to.

This class of defect can't be caught by running the benchmark. Running it is what makes it look fine. The only question that finds it is whether the input ships — which is why it's a script and not a habit.

Also audited

The other allocator figures, re-measured today, are exact: synthetic 724,298 records at parse 1.0/msg, 56 B/msg, books 0.5/msg.

228 tests pass.

…ard the rest

docs/bench/allocator.md's live row - parse 2.0 allocs/msg at 574 B, books
0.9 - was measured on captures/live-poly.feedlog. /captures/ is gitignored
as "large, regenerated". The file has never been in the repo.

So that row worked perfectly for anyone with this working tree and could
not be reproduced by anyone else. Every command in the doc runs, the
numbers come out exactly as printed, and none of it is checkable from a
clone. It is the only figure in docs/bench/ in that state, against a
README rule that a figure is either measured on a committed capture or
labelled as not yet measured. It was neither.

The capture is 85 KB gzipped and now sits beside the other six in
docs/bench/, with the gunzip-and-replay command in the doc. It reproduces
the row exactly: 877 records, parse 2.0/msg 574 B/msg, books 0.9/msg.

scripts/check_bench_inputs.py runs in CI and fails if any doc under
docs/bench/ names a .feedlog that git does not track. Angle-bracket
metavariables are excluded, since `basis xvenue-lead <capture.feedlog>`
names an argument rather than a file - the first version of the check
flagged exactly that and was wrong to.

This class of defect cannot be caught by running the benchmark. Running it
is what makes it look fine. The only question that finds it is whether the
input ships, which is why this is a script and not a habit.

Audited the other allocator figures against fresh runs while here, and
they are exact: synthetic 724,298 records at parse 1.0/msg, 56 B/msg,
books 0.5/msg.
@DanielWLiu07
DanielWLiu07 merged commit 10e3020 into main Sep 2, 2026
9 checks passed
@DanielWLiu07
DanielWLiu07 deleted the feat/ship-the-benchmark-inputs branch September 2, 2026 01:00
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant