Repository navigation
Account for accelerated HNSW host input memory - #13
Merged
nvzm123 merged 2 commits intoOct 1, 2026
Merged
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Prototype follow-up to NVIDIA/cuvs#2476, stacked on its
6c175502head. This addresses the host-memory accounting review question.ramBytesUsed()throughout population, index building, graph writing, and cleanup.InfoStream; preserve single-owner cleanup and exception propagation.Scope
This is payload accounting, not process-memory measurement or an allocation limit. It excludes upper-layer inputs, adjacency matrices, GPU workspace, and other temporary storage. Lucene's public
IndexWriter.ramBytesUsed()still uses cached buffering counters and does not automatically observe these flush/merge allocations. Ending an accounting scope does not prove successful native deallocation if cleanup fails.No public signatures, index formats, graph algorithms, flush policies, or segment-size limits change.
Validation
Local NVIDIA A10G, JDK 22.0.1, Maven 3.9.16, Lucene 10.2.0, cuVS 26.12.0:
13A: 37 passed, no skips.mvn spotless:check -Dtests.seed=13B clean verify: 333 passed, 30 skipped, no failures or errors. All six new GPU-required integration cases executed; CPU fallback fails those tests.git diff --checkpasses.The 3 GiB arithmetic/lifecycle test uses a fake allocation; this prototype has not been validated with a real multi-GiB matrix. The full suite retains existing skips and emits native-access startup, small-dataset graph-clamping, intentional invalid-workspace, and Javadoc-plugin parameter warnings.
Validation commands and prerequisite provenance
From
java/cuvs-lucene, with the Java/native prerequisites configured:mvn -o -B -Dmaven.repo.local=/tmp/pr13-m2-20261001 \ -Dtest=TestHostInputMemory,TestAcceleratedHNSWHostInputMemory,TestAcceleratedHNSWMergeReplay,TestUtilsThrowableHandling,TestAcceleratedHNSWUpperLayers,TestAcceleratedHNSWDeletedDocuments,TestAcceleratedHNSWMultiLayerRoundTrip \ -Dtests.seed=13A test mvn -o -B -Dmaven.repo.local=/tmp/pr13-m2-20261001 \ spotless:check -Dtests.seed=13B clean verifyThe isolated Maven repository reused a
cuvs-java:26.12.0JAR whose prerequisite source trees match the PR NVIDIA#2476 base. Native cuVS 26.12.0 was the existing source build at196f888f08f58ae80850d0ea8651070990ffda3c; it and cuvs-java were not rebuilt for this Java-only prototype. The cuvs-lucene module and its artifacts were rebuilt from this branch.From the repository root:
python3 fern/scripts/generate_api_reference.py --quiet(repeat run unchanged).