Skip to content

feat(api): the request half of BLPAPI, and whether short captures can be extrapolated - #67

Merged
DanielWLiu07 merged 9 commits into
mainfrom
feat/blpapi-request-half
Sep 5, 2026
Merged

DanielWLiu07 merged 9 commits into
mainfrom
feat/blpapi-request-half

Conversation

@DanielWLiu07

Copy link
Copy Markdown
Owner

Nine commits, aimed at one question: what would a market-data reviewer look for that this repo didn't have?

1. The request half of BLPAPI

The consumer interface had subscription and not request. BLPAPI has both, and the missing half is the one a screen needs when it opens — push tells you what changed, pull tells you what's true now. With only push, a late joiner waits for the next tick or subscribes to something it doesn't want.

Two decisions worth naming:

  • Entitlements are enforced identically on both. A pull path that skipped the check would be a way around it, which is how licensed data leaks in practice. The check isn't on the interesting path; it's on every path.
  • A refusal is indistinguishable from a topic that doesn't exist. "That exists but you may not see it" is itself an answer the caller wasn't entitled to, and telling them apart lets someone enumerate the topic space one probe at a time. Both outcomes are counted, so a denial stays auditable even though the caller can't tell which it was.

Validated by removing the entitlement check from the pull path — three tests fail.

2. Can a short capture stand in for a long one?

scripts/extrapolate.py slices the committed four-hour soak into windows, replays each through the real pipeline, and compares every window against the whole.

statistic whole worst window error
latency p50 0.79 µs 0.75–0.83 5.3%
latency p99 76.46 µs 1.21 98.4%
latency max 1,189 µs 104 91.3%

The median extrapolates. The tail does not — one ten-minute window reported a p99 of 1.21 µs against a true 76.46, off by 63×, in the direction that flatters.

The reason is population: the p50 describes typical messages and every window is full of those. The p99 describes rare ones, and here the rare ones are full book snapshots — a 1.13 MB Coinbase image is four orders of magnitude larger than a delta. A short window may contain none, and then its p99 is a real number measuring the wrong population.

Same family as the coordinated-omission result: a measurement that skips the expensive events reports the cheap ones and calls the result a percentile.

3. Concurrency

request() nests subscriber-inside-registry (the order publish() uses); drain() takes them sequentially so handler code can't stall publishers. Three lock disciplines in one class is where an inversion hides, and an inversion is a hang, not a wrong answer — so a single-threaded test can't find it. Added a publisher/drainer/requester test over 20,000 iterations. Clean under TSan and ASan/UBSan.

4. Documentation aimed at the actual reader

  • README now opens on what the engine is — feed handlers, normalization, distribution, entitlements — rather than on what it found. A market-data reader previously reached line 230 before learning there's a subscription API and line 420 before learning it enforces entitlements.
  • The half-open connection story moved out of a header comment: Coinbase went silent 19 minutes into a 45-minute capture and stayed silent for 25 more with zero reconnects logged, because a vanished peer produces no read error. Beast's own timeouts don't cover it.
  • Named the reference-data / market-data distinction, including that they fail differently — a stale quote announces itself, stale reference data doesn't.
  • Indexed docs/bench/ by the question each doc answers.

234 tests, perf gate, README-flag and bench-input checks all pass.

The consumer interface had subscription and not request. BLPAPI has both,
and the missing half is the one a screen needs when it opens: push tells
you what changed, pull tells you what is true now. With only push, a late
joiner either waits for the next tick or subscribes to something it does
not want to keep.

request(subscriber, event, field) returns the current value of a topic
without subscribing to it, reading the same last_value_ image a late
joiner is already seeded from.

Entitlements are enforced on it identically to publish and drain, and that
is the point rather than a detail. A pull path that skipped the check
would be a way around it, which is how licensed data leaks in practice -
the check is not on the interesting path, it is on every path.

Two decisions worth naming. A refusal is deliberately indistinguishable
from a topic that does not exist, because "that exists but you may not see
it" is itself an answer the caller was not entitled to, and telling them
apart lets a caller enumerate the topic space one probe at a time. And
both outcomes are counted, so a denial is auditable even though the caller
cannot tell which kind it was.

Five tests: the value is current rather than first, pull does not disturb
push, entitlements apply, a denial is indistinguishable from a miss, and
revocation reaches this path immediately. Validated by removing the
entitlement check from the pull path, which fails three of them.

20 -> 25 session tests, 228 -> 233 overall.
The tempting shortcut is to run a benchmark briefly and scale the result.
scripts/extrapolate.py measures whether that is legitimate, on data where
the answer is checkable: it slices the committed four-hour soak into
windows, replays each through the real pipeline, and compares every
window's estimate against the whole capture's measured value.

The median extrapolates and the tail does not. Ten-minute windows predict
the four-hour p50 within 5.3%. The same window length produced a p99 of
1.21 microseconds against a true 76.46 - off by a factor of 63, in the
direction that flatters.

The reason is the population each statistic describes. The p50 is about
typical messages and every window is full of those. The p99 is about rare
ones, and here the rare ones are full book snapshots: a 1.13 MB Coinbase
image is four orders of magnitude larger than a delta. A ten-minute window
may contain none, and when it does not, the p99 it reports is the p99 of
ordinary deltas - a real number measuring the wrong population, with
nothing in the output saying which case you got.

Same family as the coordinated omission result: a measurement that skips
the expensive events reports the cheap ones and calls the result a
percentile. There the expensive event was waiting for a turn; here it is a
snapshot arriving.

Also surfaces something about the soak itself. The venue's average rate
over the whole capture is 4.67 messages/sec against 12 to 18 over the
windows that carried traffic, because twelve of twenty-one windows held
almost nothing - the session has three real disconnects and the reconnect
gaps around them. Both figures are true and they describe different
things, which is worth stating rather than picking one.

The rule this establishes: from a short run this repo may quote the median
and ratios of counts, and may not quote any higher percentile, any
maximum, or any throughput.
The README opened on the cross-venue lead result. That is the most
interesting finding in the repo and it is the wrong first paragraph for
most readers, because it describes an outcome rather than a system - a
reader arriving from a market-data context had to reach line 230 before
learning there is a subscription API at all, and line 420 before learning
it enforces entitlements.

The opening now names the four pieces: feed handlers with gap detection
and re-snapshot, normalization onto venue-neutral event ids, conflated
distribution bounded by subscribers times topics, and default-deny
entitlements enforced on both push and pull. The lead-lag result follows,
which is the right order: what it is, then what it found.

No claim changed. Everything in the new paragraph was already true and
already documented further down; it was simply arranged so that the
interesting research result came before the working system that produced
it.
request() nests the subscriber lock inside the registry lock, which is the
order publish() uses. drain() takes the two sequentially instead,
releasing the registry lock before running handlers so that user code of
unbounded duration cannot stall publishers behind it.

Three different lock disciplines in one class is exactly where an
inversion hides, and an inversion is a deadlock rather than a wrong
answer - which means a single-threaded test cannot find it. This runs a
publisher, a drainer and a requester against each other for 20,000
iterations and asserts the counters stay coherent, so the failure mode is
a hang the suite reports rather than a rare production stall.

Clean under ThreadSanitizer and ASan/UBSan. Also documents why request()
holds the registry lock for its whole call where drain deliberately does
not: it is three hash lookups and a copy, it has to hold the lock to read
topic_ids_ and last_value_ safely, and the asymmetry would otherwise look
like an oversight next to drain's careful scoping.

25 -> 26 session tests, 233 -> 234 overall.
… comment

The best evidence in this repo that its reliability work came from
operating something rather than from reading about it was a comment in
net/ws_client.h that nobody outside the source would ever see.

A dropped network does not necessarily produce a read error. A peer or a
middlebox that disappears without a FIN or an RST leaves a blocking read
parked forever, and since the reconnect path is driven by read errors it
never runs. A 45 minute capture has Coinbase going silent 19 minutes in
and staying silent for the remaining 25 with zero reconnects logged, while
Binance recovered twice over the same outage because its server does close
connections.

Beast's stream_base::timeout does not cover it - those settings apply to
asynchronous operations and this client reads synchronously, and
keep_alive_pings has no effect while idle_timeout is none, which is the
client-role default. Hence an external watchdog thread.

That paragraph is now next to the reconnect testing it explains, because
"we handle disconnects" and "we handle the disconnects that do not
announce themselves, and here is the capture where that happened" are
different claims.
The contract registry was described as a mapping. It is reference data,
and saying so places it in the vocabulary the rest of this domain uses.

The two kinds have opposite properties and the code already treats them
differently without the doc saying why. Market data streams, changes
constantly, and is worthless once superseded - which is the entire reason
the distribution layer conflates it, keeping one slot per topic rather
than a queue. Reference data is the static description of what an
instrument is: which Kalshi ticker and which Polymarket token denote the
same outcome, what a basket's members are. It loads once at startup and a
consumer cannot interpret a quote without it.

They also fail differently, which is the part worth having written down. A
stale quote announces itself - the timestamp says so. Stale reference data
does not: a resolved contract and a quiet market both produce zero
messages, and that is precisely the failure scripts/contracts.py was
written to make loud after the registry went entirely dead unnoticed.
Ten docs in this directory and no way to see what is in them without
opening each. The index lists each by the question it answers and the
committed capture it runs on, which also makes the coverage legible: four
run on real recorded venue data, three on the deterministic synthetic
session, and every one of them names an input that ships.

Also points at the two scripts that keep this directory honest.
check_bench_inputs.py fails CI when a doc names a capture git does not
track, because that failure is invisible from a working tree - running the
benchmark is what makes it look fine. extrapolate.py answers the question
underneath all of them: whether a figure measured briefly may be quoted
for longer. Here the median may and the tail may not, by a factor of 63,
which is why the percentiles in these docs come from the thirty-minute and
four-hour captures rather than a convenient slice.
A tool whose output has to be looked up somewhere else gets run and
misread. The docstring now carries the finding: the median extrapolates
from a ten-minute window within 5.3%, the 99th percentile does not, and
the worst window reported 1.21 us against a true 76.46.

Also fixes the soak row in the bench index, which named the uncompressed
file rather than the artifact that ships.
The roadmap ranked three gaps and named the request half of BLPAPI as the
highest-leverage addition for a market-data reader. It is done, and a
roadmap that only ever grows is one nobody trusts.

Two items move to closed: the request surface, and the question of whether
a short capture can stand in for a long one. The second was not on the
original list because it had not been asked yet; the answer retroactively
justifies every percentile in docs/bench/ coming from the thirty-minute
and four-hour captures rather than a convenient slice.

Kalshi, trades, and a storage layer remain, in that order. Kalshi is still
the only one that cannot be done by writing code.
@DanielWLiu07
DanielWLiu07 merged commit 0aebdd3 into main Sep 5, 2026
9 checks passed
@DanielWLiu07
DanielWLiu07 deleted the feat/blpapi-request-half branch September 5, 2026 03:47
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant