Repository navigation
Conversation
…stination set_zero, shift_right, divmod_in_place_short and the move-assign steal path left stale limbs above the limb count, which inplace_to_bit_uint reads when the inline capacity is larger than one limb. Base 8/32 from_chars OR-ed digit blocks into stale limbs of a reused heap destination. Adds the storage_for_overwrite and clear_inline_tail helpers and inline_tail tests.
add_n_tail, sub_n_tail, lshift_copy and rshift_copy, with tests. They are not wired into basic_big_int yet.
…ocator assign_value now steals an rvalue heap source whenever the destination's allocator can free it, whatever capacity the destination holds. A propagating allocator that compares unequal no longer lets the destination keep a buffer it cannot free. Tests cover pmr, stateful propagating allocators and ownership.
…ividend in divide_unsigned
…_basecase_runtime header shortcut
…d divide dispatchers
…ail/sub_n_tail add_in_place adds or subtracts in place with the tail-aware primitives, stops touching limbs once the carry or borrow dies, grows only for the carry limb or when the other operand is longer, and sets the sign and count once. add_into writes a fresh result through storage_for_overwrite with direct stores and no zero scan or one-step trim. Integer operands now go straight through add_into. Differential tests cover operand sizes around the inline capacity, carry and borrow ripples, aliasing and constant evaluation.
lshift_copy and rshift_copy and the span_primitives and front-end test helpers cast between identical limb types, which -Wuseless-cast rejects. Use a to_limb helper in the tests, apply clang-format, and use the library's int128 aliases for the wide-integer operand test.
…py pass shift_left and shift_right work in one in-place pass and keep the inline tail zero; c = a << s and c = a >> s build the result straight from the source into a fresh, exactly sized buffer (no copy followed by an in-place shift). Shifting left no longer reserves a spare limb the value does not need, so 1 << 127 stays inline in basic_big_int<128>, and zero stays a canonical single zero limb. Negative right shifts keep floor rounding by detecting discarded bits during the pass. Differential tests cover shifts around limb boundaries, exhaustive small values, aliasing and constant evaluation.
…cator and capacity, stack scratch for small shapes
…apes, propagating allocator, constexpr
… (GCC spilled the widening_mul pair)
…shift loops under GCC Reverts 2f9150b: routing whole-limb-free in-place shifts to shift_left_n/shift_right_n slowed them on AArch64/clang by 15-50%. GCC reloads src[i - 1] in the two-load loop form when dst may alias src, so under GCC the copying shifts carry it in a register; clang keeps the two-load form, which it compiles better.
…ests for the item 1 baseline shape_sweep: sweep_int = basic_big_int<BEMAN_BIG_INT_SWEEP_INLINE_BITS> (CMake cache option, default 64) behind every timed row, printed with its sizeof in the #const line; a floor row (kernel plus one allocate/deallocate pair of the result size) for add, sub, shl, shr, mul, sqr and divrem; a vecsort op (copy + sort + sum of 100000 one-limb values, ns per element, with builtin and copy rows). gap_sweep.sh, selftest.sh and summarize.py know the new rows (auto - floor and inplace - kernel columns). Tests: move the stateful counting_allocator into util/util_counting_allocator.hpp and add alloc_count.test.cpp (steady-state allocations per call for the counting allocator, pmr default resource and basic_big_int<256>); cases that exceed their target at cf1cf54 are skipped behind item1_pending.
…n the tools README
…, seeds 1 and 2, same-seed repeat, large-band spot checks)
…library, g++-14 -march=native)
…nded; do not count the div_rem_to_zero check's own allocations
…inline-capacity study, results document and table script
…s diagnosis and fix, in-place shift A/B, correctness matrix; add regression and shift A/B scripts
…4, final regression list and open items
…cord the inline-capacity decision
C4459: rename the add_in_place/add_into local span alias (span_t -> limb_view) so it cannot hide a user's global, and the span_primitives test's max_limb, which hid a wide_ops local. C4146: no unary minus on uint64_t in muldiv_frontend. C4127 (VS 2022): if constexpr for the inline-capacity check in alloc_count.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Implements item 1 of
tests/beman/big_int/perf/gmp_gap_analysis.md: cut the fixed per-call front-end cost of+ - << >>, small*, and small division. The public API is unchanged andbig_intkeeps its one-limb inline storage. On x64, small-operand calls are 10-35% faster. On M4, shifts, mul and divare 11-39% faster, while small add/sub is flat (within about 1 ns either way). Two bugs that produced wrong values are fixed along the way.
Changes
c = a op bno longer copies the result limbs intocand frees the temporary's buffer. Allocator rules are respected: the buffer is stolen only when the destination's allocator can free it (propagating, always-equal, or comparing equal at runtime).storage_for_overwrite(n)allocates without copying or zeroing. Debug builds poison-fill new blocks; constant evaluation still constructs every limb.+=,-=andc = a +/- brun one pass through newadd_n_tail/sub_n_tailprimitives that stop once the carry dies. Signs are known without a zero scan, the result is trimmed once, and single-limb operands take a fast path.c = a << sandc = a >> sbuild the result from the source in one pass (lshift_copy/rshift_copy) into an exactly sized buffer.<<=and>>=shift in place in one pass and no longer reserve a spare limb they don't need.multiply_intowrites directly into the result.*=keeps the destination's capacity and allocator: a one-limb multiplier runs in place, products up to 64 limbs go through a stack buffer, larger ones into a fresh buffer. A header-side basecase entry,multiply_basecase_runtime, skipsthe scratch hooks and the tier ladder when
min(la, lb) < 32. It can be disabled withmul_header_basecase_enabled./=and%=keep the capacity and allocator, and divide in place in the schoolbook band.mul_add_single_limb_in_placekeeps the 128-bit product in registers. GCC spilled it to the stack on every limb; decimalfrom_charsat 2000 limbs is 10% faster on x64.Bug fixes
Base 8/32
from_charsinto a reused heapbig_intcombined new digits with stale high limbs and returned a wrong value.When the inline capacity is above one limb (
basic_big_int<128>and wider),set_zero,>>=,%=anddiv_rem_to_zero(small, large)left stale inline limbs, so==against built-in integers could be wrong.Copy/move assignment with a propagating, unequal allocator kept a heap buffer it could no longer free with the right allocator.
Behavior change
A move-assigned destination now takes the source's heap buffer even when it already has enough capacity of its own, matching
std::vector.Allocation.MoveAssignReusesDstStorageWhenLargeris inverted toMoveAssignStealsHeapSrcEvenWhenDstLarger.Results (ns, median of 4 interleaved runs, before -> after)
x64 = i9-11900K, g++-14
-march=native, pinned. M4 = Apple M4 Max, AppleClang.c = a + b, 1 limbc = a + b, 16 limbsc = a + b, 1024 limbsc = a << 13, 16 limbsc = a << 13, 2000 limbsc = a * b, 2x2c = a * b, 16x16c = a + b, 16384 limbsAllocations per call:
c *= b1 -> 0,c /= b3 -> 0,c %= b2 -> 0,a / banda % b2 -> 1,div_rem_to_zero4 -> 2 (8x4 limbs,big_int). Medium and large multiply/divide/conversion rows are within about 3%.Known gaps (follow-ups for item 2)
c = a + bat 2 and 16 limbs is about 1 ns slower (the 16-limb row is in the table above). A few in-place add/sub rows on both machines are 0.2-0.6 ns slower.c = a + bstill costs 4-6 ns beyond its unavoidable malloc/free on small operands, and divrem 8x4 costs 15-26 ns beyond its allocations.Testing
inline_tailspan_primitivesfront_end_add_subandfront_end_shift: checked against a limb-level referencemuldiv_frontenddispatch_contract: poisoned result buffers across every multiply/divide cutoffalloc_count: allocations per operation across five allocator/inline-capacity setupsallocation.test.cpp.tests/beman/big_int/perf/item1_frontend_results.md.