Repository navigation
Sum checksum16 in 64-byte blocks with an add-with-carry chain - #224
Merged
Merged
Conversation
`checksum16` read one 16-bit word per iteration into a `UInt32`. At `-Osize`, which does not vectorize the loop, that is about 2 instructions per byte, and on a link without checksum offload every UDP datagram pays it on send and receive. The accumulator also dropped carries above about 128 KiB. It now sums 64-byte blocks with a chain of `UInt128` additions, which compiles to an add-with-carry chain and needs no SIMD types, with overlapping loads for the tail and no rebinding of the caller's memory. Added tests against a word-by-word sum at every length up to 300 and every start offset up to 15, and for carries in large buffers. Measured on my Mac at `-Osize`, instructions retired per call: 85 bytes 184 -> 47 256 bytes 525 -> 81 1,200 bytes 2,413 -> 301 1,500 bytes 3,013 -> 380 At `-O` the old loop is auto-vectorized, and on a Mac it is faster than this one for buffers above about 576 bytes: about 15 ns against 40 ns at 1,500 bytes.
rnro
force-pushed
the
perf/checksum16-wide
branch
from
October 6, 2026 20:29
61d873e to
b7b4ba4
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
checksum16read one 16-bit word per iteration into aUInt32. At-Osize, which does not vectorize the loop, that is about 2 instructions per byte, and on a link without checksum offload every UDP datagram pays it on send and receive. The accumulator also dropped carries above about 128 KiB.It now sums 64-byte blocks with a chain of
UInt128additions, which compiles to an add-with-carry chain and needs no SIMD types, with overlapping loads for the tail and no rebinding of the caller's memory. Added tests against a word-by-word sum at every length up to 300 and every start offset up to 15, and for carries in large buffers.Measured on my Mac at
-Osize, instructions retired per call:85 bytes 184 -> 47
256 bytes 525 -> 81
1,200 bytes 2,413 -> 301
1,500 bytes 3,013 -> 380
At
-Othe old loop is auto-vectorized, and on a Mac it is faster than this one for buffers above about 576 bytes: about 15 ns against 40 ns at 1,500 bytes.