Skip to content

Harden parallel HNSW graph processing - #9

Merged
nvzm123 merged 2 commits into
post-ingest-hnsw-parallelismfrom
zackm_pr2653_review_fixes
Sep 30, 2026
Merged

nvzm123 merged 2 commits into
post-ingest-hnsw-parallelismfrom
zackm_pr2653_review_fixes

Conversation

@nvzm123

@nvzm123 nvzm123 commented Sep 30, 2026 •

Copy link
Copy Markdown
Owner

Important

This draft stacks the review follow-ups for NVIDIA/cuvs#2653 on its existing head branch. Keeping post-ingest-hnsw-parallelism as the base preserves the upstream PR's review history. This supersedes fork-only staging PR #8.

Base and scope

  • Base branch: nvzm123:post-ingest-hnsw-parallelism
  • Validated base: e979614fa1bd727d0cf733defe483d5c9d53c51d
  • Head branch: nvzm123:zackm_pr2653_review_fixes
  • Validated head: 03d28e150d765c66bb7a395dc64482ba93d438ff

The base includes the current PR NVIDIA#2476 head (6c175502a) and NVIDIA/cuvs main through ee29a67c9. The restack preserves PR NVIDIA#2476's ownership and cleanup fixes and main's CAGRA parameter-boundary validation.

Summary

This PR hardens the two CPU-side stages after cuVS builds an accelerated-HNSW graph:

  • Adds independent, opt-in graphThreads concurrency for graph materialization and serialization. Native cuVS concurrency remains controlled by writerThreads; both default to one for accelerated HNSW.
  • Uses one shared, bounded executor per classloader with direct handoff, caller backpressure, idle worker expiry, failure propagation, and a shared helper-thread bound.
  • Admits device-to-host adjacency copies using free physical memory and estimated native/JVM graph costs, with aggregate reservations across concurrent operations. Unsafe or unmeasurable cases remain serial.
  • Bounds serialization waves by 64 MiB of worst-case encoded data while retaining the historical 1,048,576-node absolute guardrail. Output order, bytes, and offsets remain compatible with the serial path.
  • Reports the selected graph-processing path through Lucene InfoStream.
  • Makes GPU CI fail clearly when the persisted-index sentinel cannot use cuVS, while ordinary non-GPU environments retain skip behavior.

The established public serial constructor and writeGraph(...) entry point remain available. New threaded overloads are package-private.

Coverage

The replacement tests cover:

  • writerThreads/graphThreads independence and current CAGRA boundary validation
  • serial/parallel graph and byte equivalence
  • byte- and node-bounded serialization waves
  • missing adjacency, overflow, copy failure, cleanup, and suppressed failures
  • adaptive memory estimates and aggregate reservations
  • scheduler saturation, concurrent callers, interruption, failure, and worker lifecycle
  • persisted, searchable indexes above the parallel threshold for float, binary-quantized, and scalar-quantized writers

The three earlier TestWriterThreads* classes were replaced because their premise coupled graph processing to writerThreads. Their useful coverage is retained and expanded under the independent graphThreads contract.

Validation

Final local validation used an NVIDIA A10G and the matching 26.12 Java/native stack:

  • Focused graph-processing, lifecycle, persisted-index, and parameter-boundary suite: 78 passed
  • Full mvn clean verify: 386 outcomes, 356 passed, 30 skipped, 0 failures, 0 errors
  • Java Spotless: passed
  • Fern: all 284 MDX files valid, 0 errors
  • API-reference regeneration: idempotent
  • bash -n ci/test_lucene.sh ci/test_lucene_prebuilt.sh: passed
  • git diff --check: passed

Non-failing diagnostics were the existing small-dataset graph-parameter clamp messages, JVM Vector/native-access notices, randomized-testing waits for the intentionally expiring graph workers, the Maven Javadoc-plugin warning, GDS unavailability with KvikIO fallback, and Fern's existing unauthenticated-redirect and light-mode contrast warnings.

The shared worker bound is directly tested with concurrent callers. A real concurrent multi-segment flush E2E run has not yet been performed.

Historical benchmark

At PR #8 commit 83a426b8d, a Deep1B 10M one-segment comparison reduced mean end-to-end indexing time from 36.497 s (graphThreads=1) to 28.647 s (graphThreads=16), and commit/build time from 31.509 s to 23.584 s. That benchmark predates the final observability/test changes and the latest main synchronization and was not rerun on this head.

Add independently configurable graph workers, bounded shared execution, physical-memory-aware copy admission, and byte-bounded serialization waves. Cover persistence, lifecycle, failure, concurrency, and high-degree serialization behavior.
@nvzm123
nvzm123 merged commit 03d28e1 into post-ingest-hnsw-parallelism Sep 30, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant