refactor(simd): replace std::arch and wide with pulp - #11
topi-banana wants to merge 3 commits into
Conversation
The `simd` feature carried its own vector code: a `std::arch` byte-swap for x86, scalar/NEON/wasm modules for the rest, and `wide` for the CESU-8 encode scan. Width and dispatch were settled by cfg, so the feature depended on three hand-rolled backends. Both modules now implement `pulp`'s `WithSimd` and let `Arch::dispatch` pick the widest backend the CPU has at run time. Byte-order reversal stays on 32-bit lanes, the widest integral lane pulp shifts by a runtime amount: two- and four-byte elements reverse within a lane, an eight-byte element bswaps its two lanes and then swaps them with each other. Buffers that are not four-byte aligned are staged through an aligned stack buffer, and the elements a vector block does not cover fall back to the scalar loop. pulp is taken with `default-features = false, features = ["x86-v3"]`, so the feature stays no_std and gains AVX2 dispatch on x86, NEON on aarch64 and simd128 on wasm, with pulp's scalar backend elsewhere. `simdutf8` stays for the UTF-8 check. `wide` is gone. Allocation behavior is unchanged: reading a numeric list is still one allocation (copy into the `Vec`, then swap there) and writing one is still none (the write side swaps its stack chunk before handing it to the writer). The scalar backend no longer runs on an AVX2 host, so both crates now test it directly.
fb363db to
f44e4e3
Compare
Benchmark results
Every counting platform compares its own base and pull request; a count is exact on the platform that produced it, but counts from different architectures are different instruction sets and are not comparable with each other. Platforms
linux-x86_64Callgrind instruction counts: exact within this platform, so the two sides compare without a threshold; lower is better. 60 improved · 33 regressed · 80 unchanged of 173 entries compared — 29 reference. Improved
Regressed
All 202 entries · 80 unchanged omittedparse
skip
write
linux-aarch64Callgrind instruction counts: exact within this platform, so the two sides compare without a threshold; lower is better. 40 improved · 36 regressed · 97 unchanged of 173 entries compared — 29 reference. Improved
Regressed
All 202 entries · 97 unchanged omittedparse
skip
write
wasm32-wasip1Wasmi fuel: the wasm instructions each entry runs, exact within this platform, so the two sides compare without a threshold; lower is better. 19 improved · 61 regressed · 35 unchanged of 115 entries compared — 29 reference. Improved
Regressed
All 144 entries · 35 unchanged omittedparse
skip
write
|
… libc memcpy pulp relies on #[inline(always)] on the wrappers around its intrinsics; with plain #[inline] the dispatch was not sunk into the hot loops and the rewrite lost most of what the wide-based code had. - mark the nanonbt and nanocesu8 dispatch helpers #[inline(always)] (and allow clippy::inline_always where that applies) - fuse the memcpy in decode_be: copy 8-byte chunks that are already 4-aligned directly in the reversal loop, keeping the general fallback - make the fallback copy on non-wasm targets a 64-byte read_unaligned / write_unaligned loop: for these mid-sized copies glibc's memcpy lands on rep movsb, which valgrind counts as thousands of instructions - keep copy_nonoverlapping on wasm, where the std memcpy lowers to bulk memory.copy and beats byte-wise block copies - test decode_be at every alignment and length 0..20
The compare jobs built both sides without --features simd, so every change the pulp rewrite makes sat behind #[cfg(feature = "simd")] and CI reported identical counts on both sides. Build the native compare runs with simd enabled. Do the same on wasm and execute real v128 code: add simd to WASM_FEATURES and pass -C target-feature=+simd128 to the module builds only (the host build that runs the bench needs no wasm SIMD). wasmi now runs the module with its simd feature, otherwise it refuses to load the SIMD opcodes at runtime.
Summary
Replaces the
simdfeature's hand-rolled vector code (std::archon x86, cfg-selected scalar/NEON/wasm modules, andwidefor the CESU-8 scan) withpulpbackends, written asWithSimdimplementations and dispatched at run time withArch::dispatch.Changes
nanonbt::simd: byte-order reversal moved to 32-bit lanes; two- and four-byte elements reverse within a lane, eight-byte elements bswap both lanes then swap the pair. Unaligned buffers stage through an aligned stack buffer; the elements past a vector block use the scalar loop.nanocesu8::simd: the encode fast-path scan is oneSimd-generic loop (equal_u8s/and_u8s/or_m8s+first_true_m8s);simdutf8still does UTF-8 validation.wideremoved;pulp 0.22.3added withdefault-features = false, features = ["x86-v3"](no_std; AVX2 on x86, NEON on aarch64, simd128 on wasm, pulp's scalar backend elsewhere).Behavior
Verification
cargo test --workspace --all-featuresin debug and releasecargo clippy --workspace --all-targets --all-features --locked -- -D warningsthumbv7em-none-eabiandx86_64-unknown-noneclippy + release builds withnanonbt/simd;aarch64-unknown-linux-gnuandwasm32-unknown-unknowncompile checkscargo clippy --all-targets --lockedon the pinned nightlycargo fmt --check,taplo fmt --check,cargo machete,typos