Repository navigation
feat(api): the request half of BLPAPI, and whether short captures can be extrapolated - #67
Merged
Merged
Conversation
The consumer interface had subscription and not request. BLPAPI has both, and the missing half is the one a screen needs when it opens: push tells you what changed, pull tells you what is true now. With only push, a late joiner either waits for the next tick or subscribes to something it does not want to keep. request(subscriber, event, field) returns the current value of a topic without subscribing to it, reading the same last_value_ image a late joiner is already seeded from. Entitlements are enforced on it identically to publish and drain, and that is the point rather than a detail. A pull path that skipped the check would be a way around it, which is how licensed data leaks in practice - the check is not on the interesting path, it is on every path. Two decisions worth naming. A refusal is deliberately indistinguishable from a topic that does not exist, because "that exists but you may not see it" is itself an answer the caller was not entitled to, and telling them apart lets a caller enumerate the topic space one probe at a time. And both outcomes are counted, so a denial is auditable even though the caller cannot tell which kind it was. Five tests: the value is current rather than first, pull does not disturb push, entitlements apply, a denial is indistinguishable from a miss, and revocation reaches this path immediately. Validated by removing the entitlement check from the pull path, which fails three of them. 20 -> 25 session tests, 228 -> 233 overall.
The tempting shortcut is to run a benchmark briefly and scale the result. scripts/extrapolate.py measures whether that is legitimate, on data where the answer is checkable: it slices the committed four-hour soak into windows, replays each through the real pipeline, and compares every window's estimate against the whole capture's measured value. The median extrapolates and the tail does not. Ten-minute windows predict the four-hour p50 within 5.3%. The same window length produced a p99 of 1.21 microseconds against a true 76.46 - off by a factor of 63, in the direction that flatters. The reason is the population each statistic describes. The p50 is about typical messages and every window is full of those. The p99 is about rare ones, and here the rare ones are full book snapshots: a 1.13 MB Coinbase image is four orders of magnitude larger than a delta. A ten-minute window may contain none, and when it does not, the p99 it reports is the p99 of ordinary deltas - a real number measuring the wrong population, with nothing in the output saying which case you got. Same family as the coordinated omission result: a measurement that skips the expensive events reports the cheap ones and calls the result a percentile. There the expensive event was waiting for a turn; here it is a snapshot arriving. Also surfaces something about the soak itself. The venue's average rate over the whole capture is 4.67 messages/sec against 12 to 18 over the windows that carried traffic, because twelve of twenty-one windows held almost nothing - the session has three real disconnects and the reconnect gaps around them. Both figures are true and they describe different things, which is worth stating rather than picking one. The rule this establishes: from a short run this repo may quote the median and ratios of counts, and may not quote any higher percentile, any maximum, or any throughput.
The README opened on the cross-venue lead result. That is the most interesting finding in the repo and it is the wrong first paragraph for most readers, because it describes an outcome rather than a system - a reader arriving from a market-data context had to reach line 230 before learning there is a subscription API at all, and line 420 before learning it enforces entitlements. The opening now names the four pieces: feed handlers with gap detection and re-snapshot, normalization onto venue-neutral event ids, conflated distribution bounded by subscribers times topics, and default-deny entitlements enforced on both push and pull. The lead-lag result follows, which is the right order: what it is, then what it found. No claim changed. Everything in the new paragraph was already true and already documented further down; it was simply arranged so that the interesting research result came before the working system that produced it.
request() nests the subscriber lock inside the registry lock, which is the order publish() uses. drain() takes the two sequentially instead, releasing the registry lock before running handlers so that user code of unbounded duration cannot stall publishers behind it. Three different lock disciplines in one class is exactly where an inversion hides, and an inversion is a deadlock rather than a wrong answer - which means a single-threaded test cannot find it. This runs a publisher, a drainer and a requester against each other for 20,000 iterations and asserts the counters stay coherent, so the failure mode is a hang the suite reports rather than a rare production stall. Clean under ThreadSanitizer and ASan/UBSan. Also documents why request() holds the registry lock for its whole call where drain deliberately does not: it is three hash lookups and a copy, it has to hold the lock to read topic_ids_ and last_value_ safely, and the asymmetry would otherwise look like an oversight next to drain's careful scoping. 25 -> 26 session tests, 233 -> 234 overall.
… comment The best evidence in this repo that its reliability work came from operating something rather than from reading about it was a comment in net/ws_client.h that nobody outside the source would ever see. A dropped network does not necessarily produce a read error. A peer or a middlebox that disappears without a FIN or an RST leaves a blocking read parked forever, and since the reconnect path is driven by read errors it never runs. A 45 minute capture has Coinbase going silent 19 minutes in and staying silent for the remaining 25 with zero reconnects logged, while Binance recovered twice over the same outage because its server does close connections. Beast's stream_base::timeout does not cover it - those settings apply to asynchronous operations and this client reads synchronously, and keep_alive_pings has no effect while idle_timeout is none, which is the client-role default. Hence an external watchdog thread. That paragraph is now next to the reconnect testing it explains, because "we handle disconnects" and "we handle the disconnects that do not announce themselves, and here is the capture where that happened" are different claims.
The contract registry was described as a mapping. It is reference data, and saying so places it in the vocabulary the rest of this domain uses. The two kinds have opposite properties and the code already treats them differently without the doc saying why. Market data streams, changes constantly, and is worthless once superseded - which is the entire reason the distribution layer conflates it, keeping one slot per topic rather than a queue. Reference data is the static description of what an instrument is: which Kalshi ticker and which Polymarket token denote the same outcome, what a basket's members are. It loads once at startup and a consumer cannot interpret a quote without it. They also fail differently, which is the part worth having written down. A stale quote announces itself - the timestamp says so. Stale reference data does not: a resolved contract and a quiet market both produce zero messages, and that is precisely the failure scripts/contracts.py was written to make loud after the registry went entirely dead unnoticed.
Ten docs in this directory and no way to see what is in them without opening each. The index lists each by the question it answers and the committed capture it runs on, which also makes the coverage legible: four run on real recorded venue data, three on the deterministic synthetic session, and every one of them names an input that ships. Also points at the two scripts that keep this directory honest. check_bench_inputs.py fails CI when a doc names a capture git does not track, because that failure is invisible from a working tree - running the benchmark is what makes it look fine. extrapolate.py answers the question underneath all of them: whether a figure measured briefly may be quoted for longer. Here the median may and the tail may not, by a factor of 63, which is why the percentiles in these docs come from the thirty-minute and four-hour captures rather than a convenient slice.
A tool whose output has to be looked up somewhere else gets run and misread. The docstring now carries the finding: the median extrapolates from a ten-minute window within 5.3%, the 99th percentile does not, and the worst window reported 1.21 us against a true 76.46. Also fixes the soak row in the bench index, which named the uncompressed file rather than the artifact that ships.
The roadmap ranked three gaps and named the request half of BLPAPI as the highest-leverage addition for a market-data reader. It is done, and a roadmap that only ever grows is one nobody trusts. Two items move to closed: the request surface, and the question of whether a short capture can stand in for a long one. The second was not on the original list because it had not been asked yet; the answer retroactively justifies every percentile in docs/bench/ coming from the thirty-minute and four-hour captures rather than a convenient slice. Kalshi, trades, and a storage layer remain, in that order. Kalshi is still the only one that cannot be done by writing code.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Nine commits, aimed at one question: what would a market-data reviewer look for that this repo didn't have?
1. The request half of BLPAPI
The consumer interface had subscription and not request. BLPAPI has both, and the missing half is the one a screen needs when it opens — push tells you what changed, pull tells you what's true now. With only push, a late joiner waits for the next tick or subscribes to something it doesn't want.
Two decisions worth naming:
Validated by removing the entitlement check from the pull path — three tests fail.
2. Can a short capture stand in for a long one?
scripts/extrapolate.pyslices the committed four-hour soak into windows, replays each through the real pipeline, and compares every window against the whole.The median extrapolates. The tail does not — one ten-minute window reported a p99 of 1.21 µs against a true 76.46, off by 63×, in the direction that flatters.
The reason is population: the p50 describes typical messages and every window is full of those. The p99 describes rare ones, and here the rare ones are full book snapshots — a 1.13 MB Coinbase image is four orders of magnitude larger than a delta. A short window may contain none, and then its p99 is a real number measuring the wrong population.
Same family as the coordinated-omission result: a measurement that skips the expensive events reports the cheap ones and calls the result a percentile.
3. Concurrency
request()nests subscriber-inside-registry (the orderpublish()uses);drain()takes them sequentially so handler code can't stall publishers. Three lock disciplines in one class is where an inversion hides, and an inversion is a hang, not a wrong answer — so a single-threaded test can't find it. Added a publisher/drainer/requester test over 20,000 iterations. Clean under TSan and ASan/UBSan.4. Documentation aimed at the actual reader
docs/bench/by the question each doc answers.234 tests, perf gate, README-flag and bench-input checks all pass.