Skip to content

Account for accelerated HNSW host input memory - #13

Merged
nvzm123 merged 2 commits into
zackm_cuvslucene-139from
zackm_lucene_host_memory_accounting
Oct 1, 2026
Merged

nvzm123 merged 2 commits into
zackm_cuvslucene-139from
zackm_lucene_host_memory_accounting

Conversation

@nvzm123

@nvzm123 nvzm123 commented Oct 1, 2026

Copy link
Copy Markdown
Owner

Summary

Prototype follow-up to NVIDIA/cuvs#2476, stacked on its 6c175502 head. This addresses the host-memory accounting review question.

  • Include the primary compact native host input in the float, scalar-quantized, and binary-quantized HNSW writers' ramBytesUsed() throughout population, index building, graph writing, and cleanup.
  • Report each segment/field allocation through Lucene's existing InfoStream; preserve single-owner cleanup and exception propagation.
  • Add lifecycle/error-path tests and GPU-required flush/merge coverage, including multiple fields, singleton inputs, and partial-byte binary dimensions. Update the README and regenerate affected API pages.

Scope

This is payload accounting, not process-memory measurement or an allocation limit. It excludes upper-layer inputs, adjacency matrices, GPU workspace, and other temporary storage. Lucene's public IndexWriter.ramBytesUsed() still uses cached buffering counters and does not automatically observe these flush/merge allocations. Ending an accounting scope does not prove successful native deallocation if cleanup fails.

No public signatures, index formats, graph algorithms, flush policies, or segment-size limits change.

Validation

Local NVIDIA A10G, JDK 22.0.1, Maven 3.9.16, Lucene 10.2.0, cuVS 26.12.0:

  • Focused accounting/lifecycle/deletion/multilayer run, seed 13A: 37 passed, no skips.
  • Full mvn spotless:check -Dtests.seed=13B clean verify: 333 passed, 30 skipped, no failures or errors. All six new GPU-required integration cases executed; CPU fallback fails those tests.
  • Generated API documentation is reproducible; thin-JAR checks confirm the helper and SPI descriptors are present and checked test/dependency payloads are absent; git diff --check passes.

The 3 GiB arithmetic/lifecycle test uses a fake allocation; this prototype has not been validated with a real multi-GiB matrix. The full suite retains existing skips and emits native-access startup, small-dataset graph-clamping, intentional invalid-workspace, and Javadoc-plugin parameter warnings.

Validation commands and prerequisite provenance

From java/cuvs-lucene, with the Java/native prerequisites configured:

mvn -o -B -Dmaven.repo.local=/tmp/pr13-m2-20261001 \
  -Dtest=TestHostInputMemory,TestAcceleratedHNSWHostInputMemory,TestAcceleratedHNSWMergeReplay,TestUtilsThrowableHandling,TestAcceleratedHNSWUpperLayers,TestAcceleratedHNSWDeletedDocuments,TestAcceleratedHNSWMultiLayerRoundTrip \
  -Dtests.seed=13A test
mvn -o -B -Dmaven.repo.local=/tmp/pr13-m2-20261001 \
  spotless:check -Dtests.seed=13B clean verify

The isolated Maven repository reused a cuvs-java:26.12.0 JAR whose prerequisite source trees match the PR NVIDIA#2476 base. Native cuVS 26.12.0 was the existing source build at 196f888f08f58ae80850d0ea8651070990ffda3c; it and cuvs-java were not rebuilt for this Java-only prototype. The cuvs-lucene module and its artifacts were rebuilt from this branch.

From the repository root: python3 fern/scripts/generate_api_reference.py --quiet (repeat run unchanged).

@nvzm123
nvzm123 merged commit f63c989 into zackm_cuvslucene-139 Oct 1, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant