Skip to content

Add Lucene build controls to cuVS Bench - #14

Closed
nvzm123 wants to merge 1 commit into
zackm_initial_pylucene_branchfrom
zackm_lucene_build_controls_on_2624
Closed

nvzm123 wants to merge 1 commit into
zackm_initial_pylucene_branchfrom
zackm_lucene_build_controls_on_2624

Conversation

@nvzm123

@nvzm123 nvzm123 commented Oct 1, 2026

Copy link
Copy Markdown
Owner

Dependency

Based on the feature branch for NVIDIA/cuvs#2624. Depends on NVIDIA/cuvs#2476 for reduced-memory accelerated-HNSW construction and merge handling. This diff does not include NVIDIA#2476's implementation; large-segment behavior needs combined validation.

Summary

  • Expose accelerated-HNSW m and beam_width through a stateless codec factory in the thin JAR. Record canonical parameters and label heuristic-derived graph degrees as such.
  • Add equal, sequential ingestion partitions and optional serial HNSW force merge. CAGRA retains directly built segments without force merge.
  • Verify Lucene's public per-thread RAM limit up to 2047 MiB. Larger values require explicit opt-in to a checked, unsupported reflective override; it changes the flush threshold, not memory capacity or out-of-core behavior.
  • Record requested and observed topology, the applied RAM-limit mechanism, and split build timings. Manifest schema 4 requires rebuilding older indexes with --build --force; accelerated-HNSW index names include canonical build parameters.
  • Add Java and pytest coverage, including opt-in Python/JCC cases for direct segments above 2 GiB.

Validation

To be completed after validation of this branch with NVIDIA#2476. The opt-in 3-GiB cases have not been rerun on this branch.

@nvzm123 nvzm123 closed this Oct 1, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant