Skip to content

Implement Item 1 from gmp_gap_analysis.md - #403

Open
mborland wants to merge 35 commits into
mainfrom
opt_1
Open

mborland wants to merge 35 commits into
mainfrom
opt_1

Conversation

@mborland

@mborland mborland commented Oct 6, 2026

Copy link
Copy Markdown
Member

Summary

Implements item 1 of tests/beman/big_int/perf/gmp_gap_analysis.md: cut the fixed per-call front-end cost of + - << >>, small *, and small division. The public API is unchanged and big_int keeps its one-limb inline storage. On x64, small-operand calls are 10-35% faster. On M4, shifts, mul and div
are 11-39% faster, while small add/sub is flat (within about 1 ns either way). Two bugs that produced wrong values are fixed along the way.

Changes

  • Move-assign steals the source buffer. c = a op b no longer copies the result limbs into c and frees the temporary's buffer. Allocator rules are respected: the buffer is stolen only when the destination's allocator can free it (propagating, always-equal, or comparing equal at runtime).
  • No dead zero-fills.
    • New storage_for_overwrite(n) allocates without copying or zeroing. Debug builds poison-fill new blocks; constant evaluation still constructs every limb.
    • The multiply/divide dispatchers no longer need a pre-zeroed result buffer. The tiers that accumulate now zero their own output, which removes the double fill from basecase multiplies.
  • add/sub. +=, -= and c = a +/- b run one pass through new add_n_tail/sub_n_tail primitives that stop once the carry dies. Signs are known without a zero scan, the result is trimmed once, and single-limb operands take a fast path.
  • Shifts. c = a << s and c = a >> s build the result from the source in one pass (lshift_copy/rshift_copy) into an exactly sized buffer. <<= and >>= shift in place in one pass and no longer reserve a spare limb they don't need.
  • Multiply. multiply_into writes directly into the result. *= keeps the destination's capacity and allocator: a one-limb multiplier runs in place, products up to 64 limbs go through a stack buffer, larger ones into a fresh buffer. A header-side basecase entry, multiply_basecase_runtime, skips
    the scratch hooks and the tier ladder when min(la, lb) < 32. It can be disabled with mul_header_basecase_enabled.
  • Divide. Small shapes use a 128-limb stack scratch arena instead of the heap. The remainder is built at its exact size. A 2x1 quotient that fits inline doesn't allocate. /= and %= keep the capacity and allocator, and divide in place in the schoolbook band.
  • Codegen fixes.
    • mul_add_single_limb_in_place keeps the 128-bit product in registers. GCC spilled it to the stack on every limb; decimal from_chars at 2000 limbs is 10% faster on x64.
    • Under GCC, the add/sub and shift loops use the shapes GCC compiles well: the simple carry loop and register-carried shifts. Clang and MSVC keep the unrolled and two-load forms.

Bug fixes

  • Base 8/32 from_chars into a reused heap big_int combined new digits with stale high limbs and returned a wrong value.

  • When the inline capacity is above one limb (basic_big_int<128> and wider), set_zero, >>=, %= and div_rem_to_zero(small, large) left stale inline limbs, so == against built-in integers could be wrong.

  • Copy/move assignment with a propagating, unequal allocator kept a heap buffer it could no longer free with the right allocator.

Behavior change

A move-assigned destination now takes the source's heap buffer even when it already has enough capacity of its own, matching std::vector. Allocation.MoveAssignReusesDstStorageWhenLarger is inverted to MoveAssignStealsHeapSrcEvenWhenDstLarger.

Results (ns, median of 4 interleaved runs, before -> after)

x64 = i9-11900K, g++-14 -march=native, pinned. M4 = Apple M4 Max, AppleClang.

op x64 M4
c = a + b, 1 limb 8.2 -> 5.4 6.9 -> 5.9
c = a + b, 16 limbs 25.3 -> 21.1 16.4 -> 17.7
c = a + b, 1024 limbs 842 -> 599 565 -> 450
c = a << 13, 16 limbs 24.6 -> 16.8 21.5 -> 16.1
c = a << 13, 2000 limbs 917 -> 773 654 -> 260
c = a * b, 2x2 28.2 -> 18.9 30.5 -> 18.6
c = a * b, 16x16 102 -> 89 95 -> 83
divrem 8x4 113 -> 94 96 -> 76
c = a + b, 16384 limbs -37% -23%

Allocations per call: c *= b 1 -> 0, c /= b 3 -> 0, c %= b 2 -> 0, a / b and a % b 2 -> 1, div_rem_to_zero 4 -> 2 (8x4 limbs, big_int). Medium and large multiply/divide/conversion rows are within about 3%.

Known gaps (follow-ups for item 2)

  • On x64 with GCC, in-place shifts at 256 limbs and up are 5-8% slower than before. Neither of the two loop variants measured on GCC beat the old in-place loop.
  • On M4, c = a + b at 2 and 16 limbs is about 1 ns slower (the 16-limb row is in the table above). A few in-place add/sub rows on both machines are 0.2-0.6 ns slower.
  • c = a + b still costs 4-6 ns beyond its unavoidable malloc/free on small operands, and divrem 8x4 costs 15-26 ns beyond its allocations.

Testing

  • New tests:
    • inline_tail
    • span_primitives
    • front_end_add_sub and front_end_shift: checked against a limb-level reference
    • muldiv_frontend
    • dispatch_contract: poisoned result buffers across every multiply/divide cutoff
    • alloc_count: allocations per operation across five allocator/inline-capacity setups
  • Extended allocation.test.cpp.
  • Passing configs:
    • AppleClang debug (ASan + UBSan), release and release-namespace
    • x64 clang-23 debug (ASan + UBSan) and release; g++-13 and g++-14 release; g++-14 debug; IFMA=OFF
    • Docker linux/arm64 GCC release and debug
  • The 32-bit-limb build has no new failures.
  • Full results, the inline-capacity study and the benchmark method are in tests/beman/big_int/perf/item1_frontend_results.md.

…stination

set_zero, shift_right, divmod_in_place_short and the move-assign steal path left
stale limbs above the limb count, which inplace_to_bit_uint reads when the
inline capacity is larger than one limb. Base 8/32 from_chars OR-ed digit
blocks into stale limbs of a reused heap destination. Adds the
storage_for_overwrite and clear_inline_tail helpers and inline_tail tests.
add_n_tail, sub_n_tail, lshift_copy and rshift_copy, with tests. They are not
wired into basic_big_int yet.
…ocator

assign_value now steals an rvalue heap source whenever the destination's
allocator can free it, whatever capacity the destination holds. A propagating
allocator that compares unequal no longer lets the destination keep a buffer it
cannot free. Tests cover pmr, stateful propagating allocators and ownership.
…ail/sub_n_tail

add_in_place adds or subtracts in place with the tail-aware primitives, stops
touching limbs once the carry or borrow dies, grows only for the carry limb or
when the other operand is longer, and sets the sign and count once. add_into
writes a fresh result through storage_for_overwrite with direct stores and no
zero scan or one-step trim. Integer operands now go straight through add_into.
Differential tests cover operand sizes around the inline capacity, carry and
borrow ripples, aliasing and constant evaluation.
lshift_copy and rshift_copy and the span_primitives and front-end test helpers
cast between identical limb types, which -Wuseless-cast rejects. Use a to_limb
helper in the tests, apply clang-format, and use the library's int128 aliases
for the wide-integer operand test.
…py pass

shift_left and shift_right work in one in-place pass and keep the inline tail
zero; c = a << s and c = a >> s build the result straight from the source into a
fresh, exactly sized buffer (no copy followed by an in-place shift). Shifting
left no longer reserves a spare limb the value does not need, so 1 << 127 stays
inline in basic_big_int<128>, and zero stays a canonical single zero limb.
Negative right shifts keep floor rounding by detecting discarded bits during
the pass. Differential tests cover shifts around limb boundaries, exhaustive
small values, aliasing and constant evaluation.
…cator and capacity, stack scratch for small shapes
…shift loops under GCC

Reverts 2f9150b: routing whole-limb-free in-place shifts to shift_left_n/shift_right_n
slowed them on AArch64/clang by 15-50%. GCC reloads src[i - 1] in the two-load loop
form when dst may alias src, so under GCC the copying shifts carry it in a register;
clang keeps the two-load form, which it compiles better.
…ests for the item 1 baseline

shape_sweep: sweep_int = basic_big_int<BEMAN_BIG_INT_SWEEP_INLINE_BITS> (CMake cache option, default 64) behind every
timed row, printed with its sizeof in the #const line; a floor row (kernel plus one allocate/deallocate pair of the
result size) for add, sub, shl, shr, mul, sqr and divrem; a vecsort op (copy + sort + sum of 100000 one-limb values,
ns per element, with builtin and copy rows). gap_sweep.sh, selftest.sh and summarize.py know the new rows
(auto - floor and inplace - kernel columns).

Tests: move the stateful counting_allocator into util/util_counting_allocator.hpp and add alloc_count.test.cpp
(steady-state allocations per call for the counting allocator, pmr default resource and basic_big_int<256>); cases that
exceed their target at cf1cf54 are skipped behind item1_pending.
…, seeds 1 and 2, same-seed repeat, large-band spot checks)
…nded; do not count the div_rem_to_zero check's own allocations
…inline-capacity study, results document and table script
…s diagnosis and fix, in-place shift A/B, correctness matrix; add regression and shift A/B scripts
@mborland
mborland requested a review from eisenwave as a code owner October 6, 2026 18:57
@mborland mborland self-assigned this Oct 6, 2026
@mborland mborland added the optimization Make things faster label Oct 6, 2026
C4459: rename the add_in_place/add_into local span alias (span_t -> limb_view) so it
cannot hide a user's global, and the span_primitives test's max_limb, which hid a
wide_ops local. C4146: no unary minus on uint64_t in muldiv_frontend. C4127
(VS 2022): if constexpr for the inline-capacity check in alloc_count.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

optimization Make things faster

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant