Repository navigation
feat: add brotli support for http ingress and upstream connectors - #385
Conversation
…rotli transport errors - pooledReader: deterministic test that nothing is released while a Read is parked (however many Closes arrive), exactly one release when it returns, and none afterwards - writers: a failure in Write, Flush or Close - including a connection that drops partway through a body - is cleared by the Resets between responses, checked on the same instance rather than whatever the pool hands back - brotli: a transport error after a complete stream is surfaced and stays surfaced, rather than the body being called complete - ingress: a body decoding to exactly the 32 MiB cap is served intact for gzip and br as well as zstd
Client interoperability: brotli with real HTTP clients and Ethereum librariesSetup. A local build of this branch, in front of two different public Ethereum mainnet upstreams (one connector pinned to Result. 5 languages, 204 requests, 73 Negotiation, same in every language: Request bodies. br-compressed bodies from Python Python 3.14
JavaScript (Node.js 26)
Rust 1.98
Go 1.27
Java 26
Takeaways
|
Benchmarks: gzip vs zstd vs brEnvironment. Apple M2 Pro, 12 cores, macOS 15.8, go1.27.1 darwin/arm64; this branch at Method.
Codec micro-benchmarksEncode, at production settings: gzip BestSpeed · zstd SpeedFastest with a 256 KiB window · br q1 with lgwin 18.
Decode with
Parallel encode (block, End-to-end through nodecore"Δ CPU" is nodecore's CPU-ms per 1,000 requests, minus the identity row of the same group. It is the column to compare. req/s is also shaped by moving large bodies over loopback on the same machine (see the caveats). Client-side encoding. The mock answers plain; the client sends
Upstream decoding. Block payload; the connector is pinned to
All 60 end-to-end runs had 0 errors and 0 Reading it
Caveats. The load generator, the mock and nodecore shared one laptop over loopback, with macOS scheduling across performance and efficiency cores. For large plain bodies, req/s is limited by copying bytes locally rather than by nodecore: logs with no compression reached 7.2k req/s with nodecore using only 3.6 cores. Compare codings by CPU-ms per 1k, not by req/s. CPU-ms per 1k varied by at most ±3% between repetitions. Peak RSS is sampled every 500 ms, depends on GC pacing, and is indicative only. No response caching was involved: every client request produced exactly one upstream call. |
releaseBrotliReader used to Close a decoder before resetting it into the pool. go-brrr's Close zeroes the whole decode state and drops the output buffer, so every br body - request bodies on the ingress, responses from upstreams - rebuilt its Huffman tables and buffers from scratch, and the pool saved little beyond the struct. A decoder is now only Reset, which keeps that state warm for the next body, unless its stream declared a window above lgwin 22 (4MiB): those are Closed first, so a parked reader holds at most about 8MiB - the bound a pooled zstd decoder has - and a crafted lgwin 24 stream cannot leave 32MiB in the pool. The window is read from the byte WrapReader already peeks. Decoding through WrapReader (Apple M2 Pro, mainnet bodies, lgwin <= 22) allocates 97-99% fewer bytes per body and half the allocations; a 45 B body decodes 3-4.5x faster and large bodies 4-11% faster. lgwin 24 is unchanged.
Pooled brotli decoders now stay warm across bodies (ffc0c8d)The problem
The fix
Micro-benchmarkThese decode through
Across all rows, B/op drops by 96.5% (geomean) and time by 23%. Live testSetup:
Checks:
Every valid br request body decoded to the right result. Clients and libraries:
The request bodies were encoded with five brotli encoders independent of the one nodecore uses: the reference C library (Python Web3 libraries:
Every failure the clients saw falls into one of three groups:
Negative cases:
Concurrent load1,500 requests at 48-way concurrency, paced to 30 req/s, covering all 24 combinations of request-body coding (plain, gzip, zstd, br lgwin 16 / 22 / 24; 250 each) and response coding (identity, gzip, zstd, br):
An earlier unpaced run went at about 540 req/s and hit upstream C's rate limit. Its 390 failures were all that upstream's rate-limit error passed through, with no decode or validation failures. Over the whole session, with about 3,800 client requests and about 1,750 br upstream responses decoded, the race build of nodecore logged 0 |
…g them Closing a decoder whose stream declared lgwin 23 or 24 handed its ring buffer - up to 16MiB - to go-brrr's own ring pool, where the next warm decoder that grew its ring took it and parked it, whatever its own window. Such a stream now decodes on a fresh reader that is dropped afterwards without Close, so the ring goes to the GC; it no longer takes a warm reader out of the pool either. Under a concurrent mix of crafted lgwin 24 bodies and ordinary ones, the heap still held after a GC fell from ~100MiB to under 8MiB. - the zero-allocation test reads into a fixed buffer instead of io.Copy, whose pooled buffer the race detector drops at random (8 failures in 300 -race runs before, 0 after) - the pool is used through an interface, so a test can pin that WrapReader routes each body by the window its own first byte declares; the cutoff is pinned window by window - the parked-reader memory bound is stated as it is: buffers sized to the output decoded, about 16MiB worst case on crafted input, not 8MiB
pooledReader kept its Reader and release func after releasing the codec, and a closed body stays reachable long after it is done: the ingress leaves it on the request, which echo keeps in a pooled context. For a codec that is pooled anyway that costs nothing, but a brotli decoder dropped after a large window stayed reachable through it with its ring and output buffer - 32MiB each - until the pool let the request go. The caller that takes ownership of the release now clears both under the lock; no Read can reach them again once the body is closed and idle. Measured on a live nodecore after a burst of crafted lgwin 24 bodies and ordinary ones, brotli buffers still held after the first GC went from 16-48MB to none.
Follow-up: large-window brotli decoders no longer leak into the pool (b7338f5, 87d2cb4)A second review of ffc0c8d found that its memory claim was wrong: a pooled decoder could hold much more than "about 8 MiB". Two more reviews of the fix, by different models, cross-checked it against the go-brrr source, and a live heap measurement turned up one more leak path. All of them are fixed here, and the "about 8 MiB" in my previous comment is corrected below. What was wrong
Fixes
MeasurementsIn-process, through
Live nodecore. A normal (non-race) build, with pprof enabled. The burst was 10 rounds, each of 24 concurrent requests: 4 crafted lgwin 24 bodies (33 KB on the wire, 16 MiB decoded) and 20 ordinary lgwin 22 bodies (600 KiB decoded). The heap profile was taken at the first GC after the burst, with a fresh process for each run.
Decode micro-benchmark against ffc0c8d.
lgwin 24 is the price of the fix: each such body now builds a fresh reader and ring. lgwin 24 bodies are rare in practice: request bodies mostly come from the reference CLI compressing a pipe, and the upstream that answers br in these tests declares lgwin 16. Live test, round 2Setup:
What changed from round 1:
Results:
Load:
Not changedThe other findings of that review repeat the first one:
|
What
nodecore spoke gzip and zstd on both edges. This adds brotli (
br, RFC 7932) as a third coding: the client-facing ingress servesbrto clients that negotiate it and decodesbrrequest bodies, and the HTTP upstream connectors offerbrto nodes and decode whatever they answer with. gzip and zstd behave exactly as before, and there is still no configuration.The gap it closes is clients that accept
brbut notzstd— Safari (gzip, deflate, br) and a long tail of HTTP libraries — which were held to gzip. At the level chosen here brotli is both cheaper to produce than gzip and smaller on the bodies where compression matters: a 668 KBeth_getBlockByNumberwith full transactions encodes in 719 µs to 105.5 KB, against 1,359 µs and 112.5 KB for gzip.Negotiation
…, br, zstdkeep getting zstd.Negotiateis rewritten as one loop over that preference list instead of a pair of variables per coding. Semantics are unchanged — last mention wins,*only fills codings the client did not name,q=0refuses, identity wins only when ranked strictly higher — and the existing table pins that. Three rows change answer, all becausebror a wildcard now reaches a coding nodecore speaks:br, deflate→ br,zstd;q=0, *→ br,*, zstd;q=0→ br.Accept-Encoding: zstd, br, gzip. A connector that pinsAccept-Encodingin itsheaderskeeps it, as before.internal/compressionmolecule-man/go-brrrv1.1.1 (pure Go, poolable), quality 1 with a 256 KiB window (lgwin 18, the zstd encoder's window). Quality 1 costs about half of gzip BestSpeed's CPU on large bodies and comes out 6–12% smaller; quality 0 is larger than gzip there, and quality 2 costs 1.2–2.4× the CPU for under 1%. v1.1.1 is the minimum: it is the release whoseFlushbyte-aligns below quality 2, which streamed responses rely on.WrapReaderdecodes one byte eagerly: plaintext, gzip/zstd bytes, zeroes and the non-standard large-window form are rejected up front (400 on the ingress, partial failure upstream) — the same contract gzip's header check and zstd's magic check give. An empty body, and the one-byte empty stream, pass through.io.EOF.brotliCLI declares whenever it compresses a pipe, so a lower cap would reject conformant uploads and fail conformant nodes. The cost is on crafted input: one decode can peak around 32 MiB (vs ~8 MiB for a crafted zstd frame), bounded per request; the 32 MiB decoded-body cap still applies to request bodies. This is spelled out in the compression docs.Resetand goes back to the pool warm: its tables and buffers are reused, so the next body decodes without allocating. A stream declaring lgwin 23/24 decodes on a fresh reader that is dropped afterwards and deliberately notClosed, becauseClosewould hand its up-to-16 MiB ring buffer to go-brrr's own pool for a warm reader to pick up. A parked reader's buffers are sized to what it decoded, at most about 16 MiB on crafted input. Once a body has released its codec it forgets it, so a finished request that echo keeps in its pooled context doesn't keep a dropped decoder reachable. Tests pin the cutoff, the per-body routing throughWrapReader, zero-allocation warm decodes and the release; see the comments below for measurements.Edges
The ingress middleware and the HTTP connector needed no logic changes — both reach codecs only through
Negotiate/Offer/AcquireWriter/WrapReader— so their diffs are comments, and new tests pin brotli end to end:Content-Encodingremoved; not-brotli → 400; empty accepted; past 32 MiB rejected;brhonoured; a body not in its declared coding is a partial failure; a cancelled decode is not charged to the node; a streamed response torn down while a read is in flight.One behaviour fix falls out of this:
Content-Encoding: brrequest bodies used to pass through as an unknown coding, so handlers failed to parse brotli bytes as JSON.Compatibility
No new config. Clients and nodes that don't mention
brsee no change. A client sendinggzip, deflate, brnow gets br instead of gzip. A node that speaks br and gzip but not zstd moves from gzip to br;headers: {Accept-Encoding: "zstd, gzip"}on the connector restores the old offer. gRPC and WebSocket are untouched.Testing
brotliencoder (lgwin 10, 22, 24 from a pipe, and the large-window form), so the decoder is checked against streams go-brrr did not write.make test(race, 99 packages),go vet,golangci-lintclean.zstd, br, gzip: every upstream answered zstd, unchanged;Accept-Encoding: brpinned per connector: two upstreams answered br (declared windows lgwin 22 and lgwin 16) and nodecore decoded both; the third answered plain and still worked;br,gzip, deflate, br,br, zstd,zstd,gzip, none →br,br,zstd,zstd,gzip, plain, every body decoding to the upstream's own answer for the same block (br checked with the referencebrotli -d);brotlion a pipe (lgwin 24) and from a file (lgwin 10) → 200; a plain body labelledbr→ 400;Client-library interoperability (Python, JavaScript, Rust, Go, Java and their Ethereum libraries) and gzip / zstd / br benchmarks follow in separate comments.