Skip to content

[Staging] Parallelize accelerated HNSW graph materialization and serialization - #8

Draft
nvzm123 wants to merge 2 commits into
zackm_cuvslucene-139from
staging/pr2653-review-followups-20260929
Draft

nvzm123 wants to merge 2 commits into
zackm_cuvslucene-139from
staging/pr2653-review-followups-20260929

Conversation

@nvzm123

@nvzm123 nvzm123 commented Sep 29, 2026 •

Copy link
Copy Markdown
Owner

Important

This is a fork-only staging/reference PR in nvzm123/cuvs. It provides a stable review, build, and benchmark target. It is not an upstream submission or a request to merge these changes into NVIDIA/cuvs yet.

Base and scope

This PR is an incremental change on the current PR NVIDIA#2476 branch:

  • Base branch: nvzm123:zackm_cuvslucene-139
  • Base commit: 6c175502a6d1f2b367dd4af10b63eb4091e76f7a
  • Final staging head: 04331741597b0fdc1056850f3f7f6379568f99f1
  • Scope: post-ingest accelerated-HNSW graph materialization and serialization
  • No assumptions are introduced before addDocument
  • Native cuVS graph-build concurrency remains controlled independently by writerThreads

The implementation applies to the float, scalar-quantized, and binary-quantized accelerated HNSW writers.

Summary

After cuVS builds the adjacency graph, the Java/Lucene path still has two substantial CPU stages:

  1. Materializing cuVS adjacency rows into Lucene NeighborArray instances.
  2. Encoding and serializing the Lucene level-zero graph.

This PR makes their CPU concurrency opt-in and bounded through an independent graphThreads setting. The existing serial behavior remains the default.

The latest hardening also exposes the selected materialization and serialization paths through Lucene InfoStream, verifies the real parallel path for all three accelerated writers, and makes the GPU CI lanes fail instead of silently skipping that sentinel when cuVS is unavailable.

Design

Independent graphThreads setting

AcceleratedHNSWParams gains:

.withGraphThreads(int graphThreads)
  • Valid range: 1-512
  • Default: 1
  • Includes the calling thread
  • Does not change the meaning of native writerThreads
  • Shared helper capacity can reduce effective concurrency when graph operations overlap

Shared bounded scheduler

Graph work uses one shared executor rather than creating a pool per segment or operation.

  • At most availableProcessors() - 1 helper threads, plus calling threads
  • Direct handoff with no retained work queue
  • Saturation applies backpressure by running rejected work on the caller
  • Daemon workers expire after one idle second
  • Workers do not inherit caller thread-local state or context classloaders
  • Accepted work is awaited after failure or interruption
  • Interrupt status and original failure causes are preserved

This bounds helper concurrency across concurrent segment flushes.

Parallel graph materialization

Host-backed adjacency matrices can be partitioned directly across graph workers.

Device-backed adjacency matrices require a host copy before parallel row access because their row reader is stateful. The copy is admitted only when the shared memory budget reports enough free physical memory for:

  1. The native host adjacency copy.
  2. The retained Lucene graph objects using the active JVM object layout.
  3. One additional adjacency-sized allowance for allocation races and estimation error.

Reservations are aggregated across concurrent operations in the same classloader. Invalid memory information, unsupported shapes, and arithmetic overflow fail closed to serial materialization. Temporary host allocations and reservations are released on success and failure.

Parallel graph serialization

Level-zero nodes are encoded in parallel into task-local buffers, then copied to the final IndexOutput in node order. This preserves the serial on-disk representation and offset table.

Serialization uses degree-aware bounded waves:

  • At most 1,048,576 nodes per wave
  • At most 64 MiB of worst-case encoded payload per wave
  • Higher graph degrees therefore produce smaller waves
  • Upper HNSW layers remain serial because they contain substantially less work

Parallel materialization and serialization start at 65,536 nodes.

Path observability and GPU enforcement

When the writer's Lucene InfoStream component is enabled, it reports:

  • graph-processing stage: materialization or serialization
  • selected mode: serial or parallel
  • selection reason
  • requested thread count
  • node count

The diagnostic distinguishes host-source parallel materialization, device-to-host-copy materialization, memory-admission fallback, below-threshold operation, single-thread operation, and above-threshold parallel serialization.

Both cuvs-lucene GPU CI entry points set:

CUVS_TESTS_REQUIRE_GPU=1

Under that policy, the persisted-index sentinel fails with an actionable message if cuVS cannot initialize. Ordinary environments without that opt-in retain the existing skip behavior.

Correctness and lifecycle coverage

Focused coverage includes:

  • Serial and parallel materialization equivalence
  • Byte-identical serial and parallel serialization
  • Serialization across a high-degree, byte-bounded wave boundary
  • The absolute node cap for low-degree serialization waves
  • Consistent rejection of missing adjacency data in serial and parallel serialization
  • Persisted index validation and search above the parallel threshold
  • Verified parallel materialization and serialization for float, binary-quantized, and scalar-quantized writers
  • writerThreads and graphThreads independence
  • Memory estimation at low, normal, and maximum supported graph degrees
  • Aggregate reservations across concurrent operations
  • Overflow and unavailable-memory fallback
  • Reservation release and idempotent close
  • Host-allocation cleanup when device copying fails
  • Executor saturation and caller backpressure
  • Concurrent callers sharing one worker bound
  • Failure propagation, suppressed failures, interruption, cancellation, and submission failure
  • Worker lifecycle and thread-state isolation

Validation

Final local validation ran on an NVIDIA A10G with the 26.12 Java/native stack:

  • Focused persisted-index and serialization tests: 8 passed
  • Full GPU-enabled Maven clean verify: 373 test outcomes, 0 failures, 0 errors, 30 skipped
  • Spotless: passed
  • Fern documentation: all 284 MDX files valid, 0 errors
  • API-reference regeneration: idempotent
  • bash -n ci/test_lucene.sh: passed
  • bash -n ci/test_lucene_prebuilt.sh: passed
  • git diff --check: passed

The GPU prerequisite policy was also exercised with the GPU hidden:

  • With CUVS_TESTS_REQUIRE_GPU=1, all three writer variants failed with the intended actionable initialization error.
  • Without that opt-in, the same three cases skipped and Maven succeeded.

Current validated thin JAR:

  • cuvs-lucene-26.12.0.jar: 038eac4e96a4244167ec19ffcfea0e052649cad2f6b0b76627f7f1519319bab0

The full suite still reports existing small-dataset graph-parameter clamp warnings, Vector API/native-access notices, and transient randomized-testing worker-linger notices. Fern reported two non-blocking warnings: unauthenticated redirect validation was skipped, and the existing light-mode accent contrast is below its recommended ratio.

Deep1B 10M benchmark

The performance comparison was run at staging commit 83a426b8d, before the latest diagnostic tracing, CI enforcement, invalid-adjacency check, and expanded coverage. It was not rerun on the final staging head.

Configuration:

  • Deep1B 10M, 96 dimensions, Euclidean
  • Default CAGRA heuristic
  • Graph/intermediate degree: 32/48
  • One CAGRA-HNSW layer
  • efSearch=1500, topK=1500
  • One indexing thread
  • One physical segment, no merge, no compound file
  • writerThreads=16 for both treatments
  • graphThreads=1 versus graphThreads=16
  • Cold source before every launch
  • Source and index output on local instance-store NVMe
  • 64 GiB initial / 256 GiB maximum Java heap
  • ABBA ordering with two completed runs per treatment
Configuration Indexing Commit/build fsync Ingest Recall Mean latency
Mean graphThreads=1 36.497 s 31.509 s 0.386 s 4.825 s 97.7757% 8.571 ms
Mean graphThreads=16 28.647 s 23.584 s 0.382 s 4.904 s 97.7927% 8.458 ms

Observed at that revision:

  • 21.51% lower end-to-end indexing time
  • 25.15% lower commit/build time
  • 1.274x indexing throughput
  • 1.336x commit/build throughput
  • One physical output segment in every final run
  • Peak RSS between 16.98 and 17.84 GiB

The small recall and query-latency differences are not attributed to this change. CAGRA graph construction is nondeterministic, and only 100 measured queries were used.

Two diagnostic runs accidentally used the EBS-backed default index-output location and spent approximately 35.8 seconds in fsync; they are excluded from the comparison.

Memory-policy observations

On the benchmark JVM, estimated headroom for 100 million rows was:

Graph degree Estimated headroom
1 9.2 GB
32 58.0 GB
512 826.0 GB

The host reported approximately 130.5 GB of free physical memory at measurement time. Degree 32 would therefore be eligible while degree 512 would fall back to serial materialization. These are admission-policy calculations, not 100M benchmark results.

Remaining validation

This staging snapshot has not yet established:

  • Deep1B 100M behavior
  • Jasper 10M behavior
  • Multi-segment behavior under concurrent flushes
  • High-degree performance beyond correctness at the byte-bounded wave boundary
  • Performance of the final diagnostic-enabled staging head
  • Reproduction on a second EC2 instance
  • Upstream NVIDIA CI results

Add independently configurable graph workers, bounded shared execution, physical-memory-aware copy admission, and byte-bounded serialization waves. Cover persistence, lifecycle, failure, concurrency, and high-degree serialization behavior.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant