Repository navigation
Conversation
Add independently configurable graph workers, bounded shared execution, physical-memory-aware copy admission, and byte-bounded serialization waves. Cover persistence, lifecycle, failure, concurrency, and high-degree serialization behavior.
This was referenced Sep 30, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Important
This is a fork-only staging/reference PR in
nvzm123/cuvs. It provides a stable review, build, and benchmark target. It is not an upstream submission or a request to merge these changes into NVIDIA/cuvs yet.Base and scope
This PR is an incremental change on the current PR NVIDIA#2476 branch:
nvzm123:zackm_cuvslucene-1396c175502a6d1f2b367dd4af10b63eb4091e76f7a04331741597b0fdc1056850f3f7f6379568f99f1addDocumentwriterThreadsThe implementation applies to the float, scalar-quantized, and binary-quantized accelerated HNSW writers.
Summary
After cuVS builds the adjacency graph, the Java/Lucene path still has two substantial CPU stages:
NeighborArrayinstances.This PR makes their CPU concurrency opt-in and bounded through an independent
graphThreadssetting. The existing serial behavior remains the default.The latest hardening also exposes the selected materialization and serialization paths through Lucene
InfoStream, verifies the real parallel path for all three accelerated writers, and makes the GPU CI lanes fail instead of silently skipping that sentinel when cuVS is unavailable.Design
Independent
graphThreadssettingAcceleratedHNSWParamsgains:writerThreadsShared bounded scheduler
Graph work uses one shared executor rather than creating a pool per segment or operation.
availableProcessors() - 1helper threads, plus calling threadsThis bounds helper concurrency across concurrent segment flushes.
Parallel graph materialization
Host-backed adjacency matrices can be partitioned directly across graph workers.
Device-backed adjacency matrices require a host copy before parallel row access because their row reader is stateful. The copy is admitted only when the shared memory budget reports enough free physical memory for:
Reservations are aggregated across concurrent operations in the same classloader. Invalid memory information, unsupported shapes, and arithmetic overflow fail closed to serial materialization. Temporary host allocations and reservations are released on success and failure.
Parallel graph serialization
Level-zero nodes are encoded in parallel into task-local buffers, then copied to the final
IndexOutputin node order. This preserves the serial on-disk representation and offset table.Serialization uses degree-aware bounded waves:
Parallel materialization and serialization start at 65,536 nodes.
Path observability and GPU enforcement
When the writer's Lucene
InfoStreamcomponent is enabled, it reports:The diagnostic distinguishes host-source parallel materialization, device-to-host-copy materialization, memory-admission fallback, below-threshold operation, single-thread operation, and above-threshold parallel serialization.
Both cuvs-lucene GPU CI entry points set:
Under that policy, the persisted-index sentinel fails with an actionable message if cuVS cannot initialize. Ordinary environments without that opt-in retain the existing skip behavior.
Correctness and lifecycle coverage
Focused coverage includes:
writerThreadsandgraphThreadsindependenceValidation
Final local validation ran on an NVIDIA A10G with the 26.12 Java/native stack:
clean verify: 373 test outcomes, 0 failures, 0 errors, 30 skippedbash -n ci/test_lucene.sh: passedbash -n ci/test_lucene_prebuilt.sh: passedgit diff --check: passedThe GPU prerequisite policy was also exercised with the GPU hidden:
CUVS_TESTS_REQUIRE_GPU=1, all three writer variants failed with the intended actionable initialization error.Current validated thin JAR:
cuvs-lucene-26.12.0.jar:038eac4e96a4244167ec19ffcfea0e052649cad2f6b0b76627f7f1519319bab0The full suite still reports existing small-dataset graph-parameter clamp warnings, Vector API/native-access notices, and transient randomized-testing worker-linger notices. Fern reported two non-blocking warnings: unauthenticated redirect validation was skipped, and the existing light-mode accent contrast is below its recommended ratio.
Deep1B 10M benchmark
The performance comparison was run at staging commit
83a426b8d, before the latest diagnostic tracing, CI enforcement, invalid-adjacency check, and expanded coverage. It was not rerun on the final staging head.Configuration:
efSearch=1500,topK=1500writerThreads=16for both treatmentsgraphThreads=1versusgraphThreads=16graphThreads=1graphThreads=16Observed at that revision:
The small recall and query-latency differences are not attributed to this change. CAGRA graph construction is nondeterministic, and only 100 measured queries were used.
Two diagnostic runs accidentally used the EBS-backed default index-output location and spent approximately 35.8 seconds in
fsync; they are excluded from the comparison.Memory-policy observations
On the benchmark JVM, estimated headroom for 100 million rows was:
The host reported approximately 130.5 GB of free physical memory at measurement time. Degree 32 would therefore be eligible while degree 512 would fall back to serial materialization. These are admission-policy calculations, not 100M benchmark results.
Remaining validation
This staging snapshot has not yet established: