Repository navigation
Conversation
|
Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually. Contributors can view more details about this message here. |
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configuration
📒 Files selected for processing (1)
Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 8 remain after this review. 📝 SummarySummary by CodeRabbit
WalkthroughThe stream buffer now allocates staging storage with ChangesStream Buffer Allocation
Priority: ⬇️ Low Estimated code review effort: 2 (Simple) | ~10 minutes Change: Refactor Suggested reviewers: Merge Risk: ⚪ Minimal · up to The change avoids zero-initializing the staging buffer while preserving the described write behavior. No merge-blocking issue was established; the PR is mergeable after normal checks. 🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Comment |
kvikio_ofstreamallocated its staging buffer (32 MiB by default) as astd::vector<char>, which value-initializes it, so every openmemsetthe whole buffer. With glibc a 32 MiB request is always a freshmmap, so thememsetfaults in all 8,192 pages: about 10 ms per index save, even when the file only receives a few KiB of header and the bulk data goes straight to kvikio (write_device, or host blocks at least as large as the buffer).This PR allocates the buffer with
std::make_unique_for_overwrite<char[]>and keeps its size in a member. Onlycpp/src/util/file_io.cppchanges; the public API is unchanged.[pbase(), pptr())is ever handed to kvikio, andxsputn/overflowwrite those bytes before advancingpptr(), so the bytes written are identical. A differential test of the old and newsbufproduced byte-identical files, including with the new buffer pre-filled with garbage under ASan/UBSan.Cannot open file ... for writing.A CPU microbenchmark of a 32 MiB allocation with a 4 KiB payload goes from ~10–12 ms to ~0.005 ms per open (8,193 → 2 minor page faults).
Testing
UTIL_TEST(including theFileIO.KvikioOfstream*tests) and the fullctestsuite pass.Measurements
Single-process wall time of each test executable on an RTX 6000 Ada (48 GB), otherwise idle (no ctest parallelism). The old and new
libcuvs.sowere run alternately, 2 or more repetitions each, with the order reversed between repetitions; the table shows the mean. The two builds differ only by this change and the test binaries are identical. All tests passed in every run.NEIGHBORS_ANN_VAMANA_TESTNEIGHBORS_ANN_CAGRA_FLOAT_UINT32_TESTNEIGHBORS_ANN_CAGRA_HALF_UINT32_TESTNEIGHBORS_ANN_CAGRA_INT8_UINT32_TESTNEIGHBORS_ANN_CAGRA_UINT8_UINT32_TESTNEIGHBORS_ANN_HNSW_ACE_FLOAT_UINT32_TESTNEIGHBORS_ANN_HNSW_ACE_HALF_UINT32_TESTNEIGHBORS_ANN_HNSW_ACE_INT8_UINT32_TESTNEIGHBORS_ANN_HNSW_ACE_UINT8_UINT32_TESTNEIGHBORS_ANN_CAGRA_BBQ_UINT32_TEST(not affected)NEIGHBORS_ANN_BRUTE_FORCE_TESTNEIGHBORS_ANN_IVF_SQ_TESTNEIGHBORS_ANN_IVF_RABITQ_TESTUTIL_TEST