Skip to content

feat: add brotli support for http ingress and upstream connectors - #385

Merged
mxssl merged 15 commits into
mainfrom
feat/brotli-compression
Sep 30, 2026
Merged

mxssl merged 15 commits into
mainfrom
feat/brotli-compression

Conversation

@mxssl

@mxssl mxssl commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

What

nodecore spoke gzip and zstd on both edges. This adds brotli (br, RFC 7932) as a third coding: the client-facing ingress serves br to clients that negotiate it and decodes br request bodies, and the HTTP upstream connectors offer br to nodes and decode whatever they answer with. gzip and zstd behave exactly as before, and there is still no configuration.

The gap it closes is clients that accept br but not zstd — Safari (gzip, deflate, br) and a long tail of HTTP libraries — which were held to gzip. At the level chosen here brotli is both cheaper to produce than gzip and smaller on the bodies where compression matters: a 668 KB eth_getBlockByNumber with full transactions encodes in 719 µs to 105.5 KB, against 1,359 µs and 112.5 KB for gzip.

Negotiation

  • Server preference on an exact tie is zstd > br > gzip. zstd is still the densest and by far the fastest for a client to decode; br beats gzip on both size and CPU on large bodies. Browsers that send …, br, zstd keep getting zstd.
  • Negotiate is rewritten as one loop over that preference list instead of a pair of variables per coding. Semantics are unchanged — last mention wins, * only fills codings the client did not name, q=0 refuses, identity wins only when ranked strictly higher — and the existing table pins that. Three rows change answer, all because br or a wildcard now reaches a coding nodecore speaks: br, deflate → br, zstd;q=0, * → br, *, zstd;q=0 → br.
  • The upstream offer becomes Accept-Encoding: zstd, br, gzip. A connector that pins Accept-Encoding in its headers keeps it, as before.

internal/compression

  • Encoder — molecule-man/go-brrr v1.1.1 (pure Go, poolable), quality 1 with a 256 KiB window (lgwin 18, the zstd encoder's window). Quality 1 costs about half of gzip BestSpeed's CPU on large bodies and comes out 6–12% smaller; quality 0 is larger than gzip there, and quality 2 costs 1.2–2.4× the CPU for under 1%. v1.1.1 is the minimum: it is the release whose Flush byte-aligns below quality 2, which streamed responses rely on.
  • Decoder — pooled like the others. brotli has no magic number, so WrapReader decodes one byte eagerly: plaintext, gzip/zstd bytes, zeroes and the non-standard large-window form are rejected up front (400 on the ingress, partial failure upstream) — the same contract gzip's header check and zstd's magic check give. An empty body, and the one-byte empty stream, pass through.
  • Bytes after the end of the stream are rejected. brotli has no concatenation, and gzip and zstd both reject trailing bytes. go-brrr only notices a tail that shares a buffer with the end of the stream, so the decoder is read through a small wrapper that reads one byte past io.EOF.
  • Window — the full RFC 7932 range is accepted (up to lgwin 24, 16 MiB). That is what the reference brotli CLI declares whenever it compresses a pipe, so a lower cap would reject conformant uploads and fail conformant nodes. The cost is on crafted input: one decode can peak around 32 MiB (vs ~8 MiB for a crafted zstd frame), bounded per request; the 32 MiB decoded-body cap still applies to request bodies. This is spelled out in the compression docs.
  • Decoder pooling is by declared window. A decoder whose stream declared lgwin ≤ 22 (the reference encoder's default; nodecore's own encoder uses 18) is only Reset and goes back to the pool warm: its tables and buffers are reused, so the next body decodes without allocating. A stream declaring lgwin 23/24 decodes on a fresh reader that is dropped afterwards and deliberately not Closed, because Close would hand its up-to-16 MiB ring buffer to go-brrr's own pool for a warm reader to pick up. A parked reader's buffers are sized to what it decoded, at most about 16 MiB on crafted input. Once a body has released its codec it forgets it, so a finished request that echo keeps in its pooled context doesn't keep a dropped decoder reachable. Tests pin the cutoff, the per-body routing through WrapReader, zero-allocation warm decodes and the release; see the comments below for measurements.

Edges

The ingress middleware and the HTTP connector needed no logic changes — both reach codecs only through Negotiate / Offer / AcquireWriter / WrapReader — so their diffs are comments, and new tests pin brotli end to end:

  • ingress: br served and decodable; a flushed chunk reaches a real client before the handler returns; br request bodies decoded with Content-Encoding removed; not-brotli → 400; empty accepted; past 32 MiB rejected;
  • upstream: br decoded on the buffered and streaming paths; a pinned br honoured; a body not in its declared coding is a partial failure; a cancelled decode is not charged to the node; a streamed response torn down while a read is in flight.

One behaviour fix falls out of this: Content-Encoding: br request bodies used to pass through as an unknown coding, so handlers failed to parse brotli bytes as JSON.

Compatibility

No new config. Clients and nodes that don't mention br see no change. A client sending gzip, deflate, br now gets br instead of gzip. A node that speaks br and gzip but not zstd moves from gzip to br; headers: {Accept-Encoding: "zstd, gzip"} on the connector restores the old offer. gRPC and WebSocket are untouched.

Testing

  • 14 new test functions plus brotli cases in the existing compression, ingress and connector tables, written test-first. Decoder tests use fixtures produced by the reference C brotli encoder (lgwin 10, 22, 24 from a pipe, and the large-window form), so the decoder is checked against streams go-brrr did not write.
  • make test (race, 99 packages), go vet, golangci-lint clean.
  • Live, against a local build with different public Ethereum mainnet upstreams, each behind a logging pass-through proxy:
    • default offer zstd, br, gzip: every upstream answered zstd, unchanged;
    • Accept-Encoding: br pinned per connector: two upstreams answered br (declared windows lgwin 22 and lgwin 16) and nodecore decoded both; the third answered plain and still worked;
    • client matrix br, gzip, deflate, br, br, zstd, zstd, gzip, none → br, br, zstd, zstd, gzip, plain, every body decoding to the upstream's own answer for the same block (br checked with the reference brotli -d);
    • br request bodies from brotli on a pipe (lgwin 24) and from a file (lgwin 10) → 200; a plain body labelled br → 400;
    • 1,500 requests at 48-way concurrency across all 24 combinations of response coding × request-body coding, with upstream br decoding in the mix: 0 failures, nothing logged at error level.

Client-library interoperability (Python, JavaScript, Rust, Go, Java and their Ethereum libraries) and gzip / zstd / br benchmarks follow in separate comments.

…rotli transport errors

- pooledReader: deterministic test that nothing is released while a Read is
  parked (however many Closes arrive), exactly one release when it returns,
  and none afterwards
- writers: a failure in Write, Flush or Close - including a connection that
  drops partway through a body - is cleared by the Resets between responses,
  checked on the same instance rather than whatever the pool hands back
- brotli: a transport error after a complete stream is surfaced and stays
  surfaced, rather than the body being called complete
- ingress: a body decoding to exactly the 32 MiB cap is served intact for
  gzip and br as well as zstd
@mxssl

mxssl commented Sep 23, 2026

Copy link
Copy Markdown
Contributor Author

Client interoperability: brotli with real HTTP clients and Ethereum libraries

Setup. A local build of this branch, in front of two different public Ethereum mainnet upstreams (one connector pinned to Accept-Encoding: br, so upstream br decoding was in the path as well). Every language talked to nodecore through a logging pass-through proxy that recorded the Accept-Encoding each client actually sent and the Content-Encoding it got back — the tables below come from the wire, not from what the libraries document. The big response, eth_getBlockByNumber(23000000, true) (~190 KB, 137 transactions), was validated semantically (number, hash, transaction count) in every case.

Result. 5 languages, 204 requests, 73 br responses (every one declaring lgwin 18), zero nodecore errors. Every client that can decode brotli decoded nodecore's br — through four brotli implementations independent of the one nodecore uses: the reference C library (via Python brotli/brotlicffi, Node.js zlib, Go cbrotli, Java brotli4j), Google's pure-Java decoder (org.brotli:dec, also behind OkHttp's BrotliInterceptor), rust-brotli (brotli crate, reqwest), and andybalholm/brotli (Go). Where two decoders were run on the same bytes, their output was byte-identical. The only failures were clients forced to ask for a coding they cannot decode — nodecore only sends br when it is asked for.

Negotiation, same in every language: br → br · gzip, deflate, br → br · br, zstd → zstd · zstd → zstd · gzip → gzip · no header → plain, with Vary: Accept-Encoding on every response.

Request bodies. br-compressed bodies from Python brotli, Node.js zlib, the Rust brotli crate, andybalholm/brotli (including a body flushed mid-stream) and brotli4j, at quality 5 / lgwin 22 and quality 11 / lgwin 24: single calls and 3-call JSON-RPC batches → 200 with correct results. A plain JSON body labelled Content-Encoding: br → 400 invalid compressed request body in every language.

Python 3.14

client sends by default gets br end to end
requests 2.34.2 gzip, deflate, br, zstd with brotli installed; gzip, deflate, zstd without (3.14 stdlib zstd) zstd ✅ transparent (needs brotli or brotlicffi)
httpx 0.28.1 gzip, deflate, br, zstd with extras; gzip, deflate without zstd / gzip ✅ transparent with httpx[brotli]
aiohttp 3.14.3 gzip, deflate, br, zstd / gzip, deflate, zstd zstd ✅ transparent with Brotli
urllib3 2.8.0 (bare PoolManager), urllib.request nothing plain ✅ manual decode
web3.py 8.0.0 HTTPProvider (requests) as requests zstd ✅ with brotli installed; forcing br without it → UnicodeDecodeError in the client
web3.py 8.0.0 AsyncHTTPProvider (aiohttp) as aiohttp zstd ✅ with Brotli installed; without it aiohttp refuses clearly ("Please install Brotli")

JavaScript (Node.js 26)

client sends by default gets br end to end
fetch (undici 8.11.0) gzip, deflate gzip ✅ transparent when br is requested
axios 1.20.0 gzip, compress, deflate, br br ✅ transparent
node:http + zlib, undici.request nothing plain ✅ zlib.brotliDecompressSync
viem 2.56.8 gzip, deflate gzip ✅ via http(url, { fetchOptions: { headers } })
web3.js 4.16.0 gzip,deflate gzip ✅ via provider headers
ethers 6.17.0 gzip gzip ❌ client limitation: FetchRequest always overwrites Accept-Encoding with gzip, and its Node transport only decodes gzip
ethers 5.8.0 nothing plain ❌ client limitation: a custom Accept-Encoding: br goes out, but v5 only decodes gzip and fails to parse the body

Rust 1.98

client sends by default gets br end to end
reqwest 0.12.28, no decompression features nothing plain ✅ manual decode (brotli crate)
reqwest 0.12.28, gzip + brotli + zstd zstd,gzip,br zstd ✅ transparent
reqwest 0.12.28, brotli only br br ✅ transparent
alloy 2.5.0 nothing plain ✅ once brotli is enabled on reqwest (Cargo feature unification turns it on for alloy's client), or with a reqwest client handed to connect_reqwest
ethers-rs 2.0.14 nothing plain —

Go 1.27

client sends by default gets br end to end
net/http gzip (transparent) gzip ✅ explicit header + andybalholm/brotli or cbrotli
go-ethereum 1.17.6 ethclient gzip gzip ✅ with a br-decoding RoundTripper via rpc.WithHTTPClient; rpc.WithHeader("Accept-Encoding", "br") alone → JSON error, geth ships no br decoder
lmittmann/w3 0.20.7 gzip gzip ✅ with the same transport

Java 26

client sends by default gets br end to end
java.net.http.HttpClient nothing plain ✅ manual decode (org.brotli:dec, brotli4j)
OkHttp 4.12.0 gzip gzip ✅ with BrotliInterceptor (sends br,gzip)
web3j 4.14.0 gzip gzip ✅ with BrotliInterceptor; addHeader("Accept-Encoding", "br") alone → Jackson parse error, OkHttp only decodes codings it asked for itself

Takeaways

  • Clients that offer zstd as well as br keep getting zstd, as intended. br is what reaches the clients that offer br but not zstd — axios by default, OkHttp/web3j with BrotliInterceptor, reqwest/alloy built with only the brotli feature, browsers without zstd — which is the gap this PR closes.
  • No Ethereum library tested sends br by default without an extra dependency or flag, so nothing changes for existing integrations until they opt in.
  • The failures above are all a client forcing Accept-Encoding: br without a decoder; in each case the proxy log shows nodecore answered valid br, and the same bytes decoded fine elsewhere.

@mxssl

mxssl commented Sep 23, 2026

Copy link
Copy Markdown
Contributor Author

Benchmarks: gzip vs zstd vs br

Environment. Apple M2 Pro, 12 cores, macOS 15.8, go1.27.1 darwin/arm64; this branch at 7721288, klauspost/compress v1.20.0, go-brrr v1.1.1.

Method.

  • Micro: go test ./internal/compression/ -run '^$' -bench . -benchmem -count 6 -benchtime 2s, run on the production pooled paths (AcquireWriter/ReleaseWriter, WrapReader). The figures are medians; benchstat variance was ≤ ±7%, mostly ≤ 1%.
  • End-to-end: a nodecore binary built from the same commit, in front of a local mock upstream serving real mainnet payloads. Each run used 32 keep-alive connections, a 3 s warm-up and a 15 s measured window. Mock and nodecore were restarted for every run, and each configuration ran 3 times, interleaved. The tables show the median; req/s also shows the min–max range.
  • Payloads (real mainnet responses): eth_blockNumber (45 B), eth_getBlockByNumber with transaction hashes only (26 KB, "receipt"), eth_getLogs (497 KB), and eth_getBlockByNumber with full transactions (668 KB, "block").

Codec micro-benchmarks

Encode, at production settings: gzip BestSpeed · zstd SpeedFastest with a 256 KiB window · br q1 with lgwin 18.

payload coding time/op MB/s wire bytes ratio B/op allocs/op
45 B gzip 1.38 µs 33 70 0.64 0 0
45 B zstd 145 ns 311 58 0.78 0 0
45 B br 1.95 µs 23 49 0.92 0 0
receipt 26 KB gzip 96.4 µs 271 12,921 2.02 0 0
receipt 26 KB zstd 97.3 µs 268 12,324 2.12 0 0
receipt 26 KB br 97.2 µs 269 13,690 1.91 0 0
logs 497 KB gzip 601 µs 827 50,330 9.88 0 0
logs 497 KB zstd 584 µs 851 40,730 12.21 231 2
logs 497 KB br 334 µs 1,489 44,356 11.21 276 0
block 668 KB gzip 1,334 µs 500 112,505 5.93 450 0
block 668 KB zstd 1,196 µs 558 101,030 6.61 893 2
block 668 KB br 694 µs 962 105,523 6.33 0 0

Decode with WrapReader → io.Copy(io.Discard) → Close. MB/s is measured on decoded bytes.

payload coding time/op MB/s B/op allocs/op
45 B gzip 181 ns 248 64 B 2
45 B zstd 402 ns 112 675 B 4
45 B br 2.14 µs 21 13.8 KiB 14
receipt gzip 81.6 µs 320 64 B 2
receipt zstd 29.6 µs 881 676 B 4
receipt br 102.5 µs 255 49.9 KiB 22
logs gzip 461 µs 1,078 66 B 2
logs zstd 164 µs 3,032 675 B 4
logs br 409 µs 1,217 341.5 KiB 25
block gzip 952 µs 701 76 B 2
block zstd 371 µs 1,799 676 B 4
block br 945 µs 707 379.3 KiB 25

Parallel encode (block, RunParallel, GOMAXPROCS=12): gzip 4,347 MB/s · zstd 4,919 MB/s · br 8,351 MB/s. All three pooled encoders scale about 8.7× over one thread.

End-to-end through nodecore

"Δ CPU" is nodecore's CPU-ms per 1,000 requests, minus the identity row of the same group. It is the column to compare. req/s is also shaped by moving large bodies over loopback on the same machine (see the caveats).

Client-side encoding. The mock answers plain; the client sends Accept-Encoding: <coding>.

payload client AE req/s (min–max) p50 ms p99 ms wire B/resp CPU-ms / 1k Δ CPU peak RSS MiB
45 B — 47,857 (42,781–48,111) 0.59 1.99 45 110 — 245
45 B gzip 48,538 (46,588–48,906) 0.58 2.01 70 112 +2 293
45 B zstd 48,828 (45,061–48,862) 0.58 2.00 58 109 −1 280
45 B br 47,948 (44,806–48,010) 0.59 2.00 49 114 +4 255
receipt — 29,914 (27,695–30,739) 0.92 3.36 26,128 182 — 288
receipt gzip 22,183 (20,908–22,907) 1.22 4.80 12,921 309 +127 323
receipt zstd 23,373 (21,655–23,410) 1.19 4.17 12,324 313 +131 299
receipt br 23,850 (22,140–24,070) 1.16 4.10 13,689 295 +113 305
logs — 7,223 (7,001–7,463) 4.36 8.21 497,375 497 — 200
logs gzip 6,927 (6,549–7,144) 3.95 15.94 50,330 1,080 +583 243
logs zstd 7,336 (6,852–7,622) 3.73 14.92 40,730 1,030 +533 260
logs br 10,116 (9,380–10,269) 2.73 10.24 44,356 716 +219 282
block — 4,007 (3,586–4,037) 7.43 18.70 667,656 1,563 — 246
block gzip 2,609 (2,273–2,627) 10.74 34.44 112,505 3,092 +1,529 464
block zstd 2,954 (2,863–3,023) 9.49 31.54 101,030 2,761 +1,198 476
block br 3,672 (3,603–3,744) 7.89 23.18 105,523 2,130 +567 430

Upstream decoding. Block payload; the connector is pinned to Accept-Encoding: <coding> and the mock answers in that coding, using each library's default level. The client receives plain.

upstream coding (wire B) req/s (min–max) p50 ms p99 ms CPU-ms / 1k Δ CPU peak RSS MiB
identity (667,656) 3,963 (3,828–4,022) 7.52 18.69 1,566 — 244
gzip (111,151) 3,245 (3,143–3,271) 9.00 24.34 2,620 +1,054 260
zstd (104,437) 4,860 (4,764–4,996) 5.96 18.18 1,634 +68 486
br (105,702) 3,608 (3,424–3,690) 7.73 25.29 2,379 +813 382

All 60 end-to-end runs had 0 errors and 0 Content-Encoding mismatches, and every run's first response decoded to a valid JSON-RPC result. nodecore logged no errors.

Reading it

  • Encoding large bodies: br at quality 1 is the cheapest of the three encoders. It runs at about 0.5× gzip's and 0.6× zstd's CPU on block and logs, both in the micro-benchmark and end to end (+567 against +1,198 for zstd and +1,529 for gzip CPU-ms per 1k on block; +219 against +533 / +583 on logs).
  • Bytes on large bodies: br sits between zstd and gzip. On block it is 6% smaller than gzip and 4% larger than zstd; on logs, 12% smaller than gzip and 9% larger than zstd.
  • Mid-size bodies (26 KB): all three encoders cost the same, about 97 µs. br's output there is the largest of the three (13.7 KB, against 12.9 KB for gzip and 12.3 KB for zstd).
  • Tiny bodies: there is no size threshold, so all three codings make a 45 B reply bigger (br 49 B, zstd 58 B, gzip 70 B). Per-request CPU doesn't change beyond noise. This behaviour already existed for gzip and zstd.
  • Decoding: br is the slowest decoder of the three, about equal to gzip on large bodies and roughly 2.5× slower than zstd. Upstream, a br answer costs nodecore +813 CPU-ms per 1k, against +1,054 for gzip and +68 for zstd. That's why the upstream offer keeps zstd first.
  • Allocations on br decode: each large br decode allocates about 0.4 MB, against under 1 KiB for gzip and zstd.
    • Cause: when a pooled reader is released, the branch calls go-brrr's Reader.Close so its ring buffer goes back to the library's pool, rather than keeping up to a 16 MiB window parked in every pooled reader. Close also drops the reader's output staging buffer, and the next decode grows it again.
    • This is short-lived garbage, not retained memory. It is a possible follow-up if br decoding ever shows up in profiles; go-brrr doesn't expose a way to keep one buffer and drop the other.
  • Parallel encoding: br keeps about 1.9× gzip's aggregate encode throughput across 12 threads.

Caveats. The load generator, the mock and nodecore shared one laptop over loopback, with macOS scheduling across performance and efficiency cores. For large plain bodies, req/s is limited by copying bytes locally rather than by nodecore: logs with no compression reached 7.2k req/s with nodecore using only 3.6 cores. Compare codings by CPU-ms per 1k, not by req/s. CPU-ms per 1k varied by at most ±3% between repetitions. Peak RSS is sampled every 500 ms, depends on GC pacing, and is indicative only. No response caching was involved: every client request produced exactly one upstream call.

@mxssl
mxssl requested review from a team and KirillPamPam September 24, 2026 10:52
releaseBrotliReader used to Close a decoder before resetting it into the
pool. go-brrr's Close zeroes the whole decode state and drops the output
buffer, so every br body - request bodies on the ingress, responses from
upstreams - rebuilt its Huffman tables and buffers from scratch, and the
pool saved little beyond the struct.

A decoder is now only Reset, which keeps that state warm for the next body,
unless its stream declared a window above lgwin 22 (4MiB): those are Closed
first, so a parked reader holds at most about 8MiB - the bound a pooled zstd
decoder has - and a crafted lgwin 24 stream cannot leave 32MiB in the pool.
The window is read from the byte WrapReader already peeks.

Decoding through WrapReader (Apple M2 Pro, mainnet bodies, lgwin <= 22)
allocates 97-99% fewer bytes per body and half the allocations; a 45 B body
decodes 3-4.5x faster and large bodies 4-11% faster. lgwin 24 is unchanged.
@mxssl

mxssl commented Sep 29, 2026

Copy link
Copy Markdown
Contributor Author

Pooled brotli decoders now stay warm across bodies (ffc0c8d)

The problem

releaseBrotliReader called go-brrr's Reader.Close() before Reset(nil). Close does r.state = decodeState{} and r.out = nil, which throws away the Huffman tables, arenas and output buffer that initForReuse is written to keep. So every br body rebuilt its decode state from scratch, and the pool saved almost nothing but the struct. That affected both request bodies on the ingress and responses from upstreams. It is also why br decode allocated ~0.4 MB per large body in the benchmarks above.

Close was there so a parked reader wouldn't keep a large window. But Close only moves the ring buffer into go-brrr's own sync.Pool, so the memory stays pooled until the next GC either way. go-brrr's README recommends pooling readers and calling Reset, without Close.

The fix

  • Warm by default. A released decoder is now only Reset, so its state stays warm for the next body.
  • Cold for large windows. A stream whose first byte declared a window above lgwin 22 (4 MiB) is Closed first, so it starts cold. The large-window form is too.
    • A parked reader keeps a ring buffer and an output buffer, each up to its window. That caps it at about 8 MiB, the same bound as decoderMaxWindow for a pooled zstd decoder.
    • A crafted lgwin 24 stream still can't leave 32 MiB parked in the pool.
    • lgwin 22 covers what encoders declare by default: the reference library uses 22, and nodecore's own encoder uses 18.
  • No extra read. The window comes from the byte WrapReader already peeks (RFC 7932 §9.1), so there is no extra I/O.
  • New tests:
    • A released lgwin 10 / 22 decoder decodes its next body with 0 allocations, while an lgwin 24 decoder starts cold. Always calling Close (the old behaviour) fails this test, and so does never calling it.
    • One warm decoder decodes the next body exactly after each of: a larger window, a smaller one, a stream abandoned halfway, and a body that failed to decode.
    • Every branch of the window encoding is covered.
  • No duplicated parser. The test-only window parser moved into production code, and the black-box tests use it through export_test.go.

make test (race) passes except one test, TestPostgresConnectorStoreAndRemoveExpired. That one is a known timing flake: it passes 4 of 4 when re-run on its own, and its package doesn't import internal/compression. golangci-lint reports 0 issues.

Micro-benchmark

These decode through WrapReader → io.Copy(io.Discard) → Close, measured with benchstat, 6 runs each, on an Apple M2 Pro. The bodies are real mainnet responses: eth_blockNumber (45 B), block 23000000 with tx hashes (13 KB), eth_getLogs for that block (213 KB), and the same block with full txs (184 KB). Each body is encoded three ways: q1/lgwin 18 (what nodecore's encoder produces), q5/lgwin 22 (the reference default window) and q5/lgwin 24.

body window time before → after B/op before → after allocs
45 B lgwin 18 1.84 µs → 0.41 µs (−78%) 14,138 → 389 15 → 11
45 B lgwin 22 3.67 µs → 1.14 µs (−69%) 23,620 → 389 23 → 11
13 KB lgwin 18 46.6 → 43.0 µs (−8%) 37,154 → 391 23 → 11
13 KB lgwin 22 44.3 → 40.4 µs (−9%) 39,487 → 391 23 → 11
213 KB logs lgwin 18 180 → 168 µs (−6%) 161,139 → 555 24 → 12
213 KB logs lgwin 22 182 → 163 µs (−11%) 247,121 → 426 23 → 11
184 KB block lgwin 18 324 → 311 µs (−4%) 161,114 → 662 24 → 12
184 KB block lgwin 22 361 → 345 µs (−4%) 186,279 → 533 23 → 11
any lgwin 24 unchanged (cold by design) unchanged unchanged

Across all rows, B/op drops by 96.5% (geomean) and time by 23%.

Live test

Setup:

  • nodecore was built from this branch with -race.
  • Clients reached it through a logging pass-through proxy. The tables below come from what that proxy recorded on the wire.
  • Behind nodecore were three different public Ethereum mainnet upstreams, each behind its own logging proxy. Every connector was pinned to Accept-Encoding: br, so br from upstreams went through the changed decoder too.
    • Upstream A answered br (lgwin 16).
    • Upstream B speaks only gzip/zstd, so it answered plain to a br-only offer.
    • Upstream C answered plain.

Checks:

  • Every response was checked for content, not just status. For block 0x18e0028: hash, 296 txs, and in the typed clients the block and tx hashes recomputed from the decoded fields. For eth_getLogs on that block: all 603 logs.
  • br request bodies were sent at lgwin 16 and 22 (warm path) and lgwin 24 (cold path), at quality 5 and 11, interleaved (16, 24, 22, 24, 16, 22 …) so warm and cold decoders kept alternating.
  • The payloads were single calls, 3-call batches, and 120–290 KB bodies whose back-references reach past a 64 KiB window. Some bodies were flushed mid-stream, and two were streamed chunked with a pause after the flush.
  • Every negative case was followed by valid bodies.
  • For every body, the proxy confirmed the declared window was the one intended.
language requests br request bodies (lgwin 16 / 22 / 24) br responses decoded nodecore failures
Python 3.14 264 104 (35 / 34 / 35) 122 0
JavaScript (Node.js 26) 182 92 (29 / 31 / 32) 116 0
Rust 1.98 162 86 (28 / 30 / 28) 82 0
Go 1.27 119 67 (22 / 21 / 24) 58 0
Java 27 72 42 (13 / 14 / 15) 40 0
total 799 391 (127 / 130 / 134) 418 0

Every valid br request body decoded to the right result.

Clients and libraries:

  • Python: requests, httpx, aiohttp, urllib3, urllib.request, and web3.py (HTTPProvider, AsyncHTTPProvider).
  • JavaScript: fetch, undici.request (with and without the decompress interceptor), axios, node:http + zlib, viem, ethers 6 and 5, and web3.js.
  • Rust: reqwest in three feature sets, alloy and ethers-rs.
  • Go: net/http with andybalholm/brotli, cbrotli and go-brrr, go-ethereum ethclient/rpc, and lmittmann/w3.
  • Java: java.net.http.HttpClient with org.brotli:dec and brotli4j, OkHttp with and without BrotliInterceptor, and web3j.

The request bodies were encoded with five brotli encoders independent of the one nodecore uses: the reference C library (Python brotli, Node.js zlib, cbrotli, brotli4j), the Rust brotli crate, andybalholm/brotli, and go-brrr.

Web3 libraries:

library br responses br request bodies
web3.py 8.0.0 ✅ with brotli installed ✅ via a custom requests.Session / aiohttp.ClientSession
viem 2.57.0 ✅ via fetchOptions.headers ✅ via a custom fetchFn, incl. batch
web3.js 4.16.0 ✅ via provider headers ✅ via an HttpProvider subclass, incl. BatchRequest
ethers 6.17.0 / 5.8.0 ✅ with a fetch-based transport; ❌ built-in transport (it only decodes gzip) ✅ with a custom transport
alloy 2.5.0 / ethers-rs 2.0.14 ✅ once the brotli feature is on the reqwest version the library itself uses n/a
go-ethereum 1.17.6 / w3 0.20.7 ✅ with a br-decoding RoundTripper via rpc.WithHTTPClient ✅ via the same RoundTripper
web3j 4.14.0 ✅ with OkHttp + BrotliInterceptor; ❌ with only an Accept-Encoding: br header ✅ via an OkHttp interceptor

Every failure the clients saw falls into one of three groups:

  • Client limitations. A client forced Accept-Encoding: br without a decoder, and the bytes on the wire were valid br: Python without brotli, the ethers built-in transports, alloy and ethers-rs without the reqwest feature, plain net/http / go-ethereum, and web3j with only the header.
  • One upstream. Upstream B answered HTTP 500 to a 140 KB request 3 of 3 times. The proxy log shows nodecore forwarded the correctly decoded 140,074-byte body each time, and the same body succeeded on the other upstreams.
  • A harness bug. An ethers 6 adapter reused the uncompressed content-length with a compressed body, so the client aborted 4 requests. That run was discarded and re-run after the adapter was fixed.

Negative cases:

  • A plain JSON body labelled br got 400 invalid compressed request body.
  • A br body cut off halfway also got a 400, but it's the JSON-RPC couldn't parse a request, because the stream only fails once the handler reads it. gzip and zstd give the identical response for a truncated body, so this isn't new with br.

Concurrent load

1,500 requests at 48-way concurrency, paced to 30 req/s, covering all 24 combinations of request-body coding (plain, gzip, zstd, br lgwin 16 / 22 / 24; 250 each) and response coding (identity, gzip, zstd, br):

  • 1,500 / 1,500 passed.
  • Upstream A sent 334 br responses during the run, and nodecore decoded all of them through the warm path.

An earlier unpaced run went at about 540 req/s and hit upstream C's rate limit. Its 390 failures were all that upstream's rate-limit error passed through, with no decode or validation failures.

Over the whole session, with about 3,800 client requests and about 1,750 br upstream responses decoded, the race build of nodecore logged 0 DATA RACE reports and 0 errors.

…g them

Closing a decoder whose stream declared lgwin 23 or 24 handed its ring
buffer - up to 16MiB - to go-brrr's own ring pool, where the next warm
decoder that grew its ring took it and parked it, whatever its own window.
Such a stream now decodes on a fresh reader that is dropped afterwards
without Close, so the ring goes to the GC; it no longer takes a warm reader
out of the pool either. Under a concurrent mix of crafted lgwin 24 bodies
and ordinary ones, the heap still held after a GC fell from ~100MiB to
under 8MiB.

- the zero-allocation test reads into a fixed buffer instead of io.Copy,
  whose pooled buffer the race detector drops at random (8 failures in 300
  -race runs before, 0 after)
- the pool is used through an interface, so a test can pin that WrapReader
  routes each body by the window its own first byte declares; the cutoff is
  pinned window by window
- the parked-reader memory bound is stated as it is: buffers sized to the
  output decoded, about 16MiB worst case on crafted input, not 8MiB
pooledReader kept its Reader and release func after releasing the codec,
and a closed body stays reachable long after it is done: the ingress
leaves it on the request, which echo keeps in a pooled context. For a
codec that is pooled anyway that costs nothing, but a brotli decoder
dropped after a large window stayed reachable through it with its ring
and output buffer - 32MiB each - until the pool let the request go.

The caller that takes ownership of the release now clears both under the
lock; no Read can reach them again once the body is closed and idle.
Measured on a live nodecore after a burst of crafted lgwin 24 bodies and
ordinary ones, brotli buffers still held after the first GC went from
16-48MB to none.
@mxssl

mxssl commented Sep 29, 2026

Copy link
Copy Markdown
Contributor Author

Follow-up: large-window brotli decoders no longer leak into the pool (b7338f5, 87d2cb4)

A second review of ffc0c8d found that its memory claim was wrong: a pooled decoder could hold much more than "about 8 MiB". Two more reviews of the fix, by different models, cross-checked it against the go-brrr source, and a live heap measurement turned up one more leak path. All of them are fixed here, and the "about 8 MiB" in my previous comment is corrected below.

What was wrong

  1. Close put the 16 MiB rings into circulation. Close on a decoder that had decoded an lgwin 24 stream handed its ring buffer (up to 16 MiB) to go-brrr's own ring pool. The next decoder that needed a ring took it, whatever its own window, and kept it parked warm. The code before ffc0c8d had the same problem, only in a different pool. ffc0c8d didn't fix it.
  2. Released bodies kept their codec reachable. A pooledReader still pointed at its codec after releasing it, both through Reader and through the release func. A finished body can stay reachable for a while, because the ingress leaves it on the request and echo keeps the request in its pooled context. A dropped large-window decoder therefore stayed alive, holding its ring and output buffer (16 MiB each), until that pool let the request go.
  3. The zero-allocation test was flaky under -race. It failed 8 times in 300 runs. It went through io.Copy, whose pooled buffer the race detector drops at random.
  4. The documented bound was too low. The docs said "about 8 MiB". The real worst case on crafted input is about 16 MiB, made up of:
    • the output buffer, which can grow past its 4 MiB window with append slack;
    • an up-to-8 MiB ring, recycled from a larger window's decode;
    • up to ~2.5 MiB of Huffman tables.

Fixes

  • Large windows get a fresh reader. A stream whose first byte declares lgwin 23 or 24 is decoded on a fresh reader, and the reader is dropped without Close afterwards, so its ring goes to the GC. It also no longer takes a warm reader out of the pool.
  • A released body forgets its codec. Whoever takes ownership of the release clears both Reader and release, under the lock.
  • The docs state the bound as it is.
    • A parked reader's buffers are sized to the output it decoded: for ordinary RPC bodies, each is at most the largest body rounded up to a power of two.
    • The worst case on crafted input is about 16 MiB.
    • The code comment, 15-compression.md, the spec and the PR description are updated.
  • Tests:
    • The warm decode is checked for zero allocations while reading into a fixed buffer, so it doesn't depend on sync.Pool.
    • The lgwin 22/23 cutoff is pinned window by window.
    • A counting wrapper around the pool checks that WrapReader routes each body by the window its own first byte declares. That includes the reference encoder's empty stream 0x3f, which declares lgwin 24, and the large-window form. Hard-coding the window in wrapBrotliReader fails this test, and so does moving the cutoff to 23.
    • A weak.Pointer check shows that a released body leaves its codec collectable, both on an idle Close and on a Close that races a parked Read.
    • A large-window decoder is checked to be dropped, not Closed.
  • Results: make test (race) passes all 99 packages, and golangci-lint reports 0 issues. The brotli tests passed 300 runs under -race.

Measurements

In-process, through WrapReader. 8 rounds, each of 6 concurrent bodies: two crafted lgwin 24 bodies (16 MiB decoded) and four ordinary ones. The number is the heap still held after one GC.

build heap held
ffc0c8d 96–120 MiB
this branch under 8 MiB

Live nodecore. A normal (non-race) build, with pprof enabled. The burst was 10 rounds, each of 24 concurrent requests: 4 crafted lgwin 24 bodies (33 KB on the wire, 16 MiB decoded) and 20 ordinary lgwin 22 bodies (600 KiB decoded). The heap profile was taken at the first GC after the burst, with a fresh process for each run.

build go-brrr memory still held whole heap
ffc0c8d, 3 runs 16 / 32 / 48 MB (all getDecRingBuf, i.e. 16 MiB rings) 29 / 49 / 64 MB
this branch, 3 runs 0 / 0 / 0 12–19 MB (baseline)

Decode micro-benchmark against ffc0c8d. benchstat, n=6, same mainnet bodies as before.

window time allocation per body
lgwin ≤ 22 unchanged (within ±4%) unchanged, ~390 B
lgwin 24 8–17% slower on 13–213 KB bodies; 2.4× slower on a 45 B body (3.7 → 8.8 µs) 2–3× more

lgwin 24 is the price of the fix: each such body now builds a fresh reader and ring. lgwin 24 bodies are rare in practice: request bodies mostly come from the reference CLI compressing a pipe, and the upstream that answers br in these tests declares lgwin 16.

Live test, round 2

Setup:

  • Same stack as before: nodecore built from this branch with -race, behind a logging proxy, with every connector pinned to br, and three different public mainnet upstreams behind their own logging proxies.
  • The reference block moved to 0x18e0210 (hash 0x2641…60a0, 293 txs, 876 logs), because one upstream only serves recent blocks.
  • Transactions and logs were compared with the reference field by field.

What changed from round 1:

  • lgwin 23 added. The interleaved request-body rounds gained lgwin 23, so bodies cross the 22/23 boundary back to back in both directions.
  • Empty streams. 0x3b (lgwin 22), 0x3f (lgwin 24) and a zero-length body, each labelled br, were compared with a plain empty body.
  • Far back-references. Go and Java added 9–9.4 MB bodies whose back-references reach further than 4 MiB. They compress to about half the size at lgwin 23/24 as at lgwin 22, which shows the far references are really in the stream.
language requests valid br request bodies (lgwin 16 / 22 / 23 / 24) nodecore failures
Python 3.14 250 144 (35 / 40 / 37 / 32) 0
JavaScript (Node.js 26) 199 135 (30 / 41 / 33 / 31) 0
Rust 1.98 173 113 (27 / 30 / 30 / 26) 0
Go 1.27 161 100 (21 / 30 / 25 / 24) 0
Java 27 106 68 (15 / 19 / 18 / 16) 0
total 889 560 (128 / 160 / 143 / 129) 0

Results:

  • Request bodies. Every valid br request body decoded to the right result. That covers quality 5 and 11, single calls and batches, bodies flushed mid-stream, bodies of 120 KB to 9.4 MB, and bodies sent through each web3 library's custom transport. The clients and libraries are the same as in round 1:
    • Python: requests, httpx, aiohttp, urllib3 and web3.py.
    • JavaScript: fetch, undici, axios, node:http, viem, ethers 6 and 5, and web3.js.
    • Rust: reqwest, alloy and ethers-rs.
    • Go: net/http with andybalholm/brotli, cbrotli and go-brrr, go-ethereum and w3.
    • Java: HttpClient, OkHttp and web3j.
  • Empty streams. 0x3b, 0x3f and the zero-length body got a byte-identical answer to a plain empty body in every language. So did the normal br body sent right after each one.
  • Negatives. A plain body labelled br and a truncated stream got the same 400s as in round 1. The bodies sent after them decoded fine.
  • Client limitations. The ones seen are unchanged from round 1: clients that force Accept-Encoding: br without a decoder.

Load:

  • 1,500 requests at 48-way concurrency, paced to 30 req/s, covering all 28 combinations of request-body coding (plain, gzip, zstd, br lgwin 16 / 22 / 23 / 24) × response coding: 1,500 / 1,500 passed.
  • The load test and the Go suite were run again on the final commit (87d2cb4): 1,500 / 1,500 and 146 / 146. Upstream A sent 749 br responses during that load run, and nodecore decoded them through the warm path.
  • The race build logged 0 DATA RACE reports.

Not changed

The other findings of that review repeat the first one:

  • lgwin cap at 22: not done, because it would reject conformant bodies. With this fix, a crafted lgwin 24 body's memory is freed after the request instead of being pooled.
  • Metadata blocks with no output: not specific to brotli, since zstd skippable frames and empty deflate blocks behave the same way; the limit is ReadTimeout.
  • br preferred over gzip on ties: kept, because the end-to-end numbers show br costs no more CPU than gzip at any body size.
  • br in the upstream offer: kept, with the documented per-connector opt-out.

@mxssl
mxssl merged commit 58cfd87 into main Sep 30, 2026
5 checks passed
@mxssl
mxssl deleted the feat/brotli-compression branch September 30, 2026 12:23
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants