Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 4 additions & 0 deletions fern/docs.yml
Original file line number Diff line number Diff line change
Expand Up @@ -580,6 +580,10 @@ navigation:
path: "./pages/lucene_api/lucene-api-com-nvidia-cuvs-lucene-gpusearchparams.md"
- page: "IndexSearcherTimingBridge"
path: "./pages/lucene_api/lucene-api-com-nvidia-cuvs-lucene-indexsearchertimingbridge.md"
- page: "IndexWriterConfigRAMLimitBridge"
path: "./pages/lucene_api/lucene-api-com-nvidia-cuvs-lucene-indexwriterconfigramlimitbridge.md"
- page: "Lucene101AcceleratedHNSWCodecFactory"
path: "./pages/lucene_api/lucene-api-com-nvidia-cuvs-lucene-lucene101acceleratedhnswcodecfactory.md"
- page: "LuceneProvider"
path: "./pages/lucene_api/lucene-api-com-nvidia-cuvs-lucene-luceneprovider.md"
- page: "ThreadLocalCuVSResourcesProvider"
Expand Down
185 changes: 184 additions & 1 deletion fern/pages/cuvs_bench/lucene_backend.md
Original file line number Diff line number Diff line change
Expand Up @@ -26,6 +26,113 @@ Selecting this backend is explicit. Commands that do not select it with
explicit `--algorithms` value. Select `lucene_cpu_hnsw` explicitly when a CPU
control is needed.

The accelerated-HNSW build accepts Lucene's `m` and `beam_width` parameters.
Both must be integers in the range 1 through 512 when either is specified;
omitting both preserves the codec defaults of 32 and 32. cuVS derives the
CAGRA build parameters with the `SAME_GRAPH_FOOTPRINT` heuristic, so `m: 16`
produces `graph_degree=32` and `intermediate_graph_degree=48`. For example, an
algorithm configuration for `m=16` and `beam_width=80` is:

```yaml
name: lucene_accelerated_hnsw
groups:
m16_bw80:
build:
codec: ["Lucene101AcceleratedHNSWCodec"]
m: [16]
beam_width: [80]
search: {}
```

Pass that file with `--configuration`, select
`--algorithms lucene_accelerated_hnsw --groups m16_bw80`, and use `--build`.
The normalized HNSW parameters are recorded in the index manifest. Result
metadata records those requested parameters together with graph degrees derived
from the selected heuristic; the graph-degree fields are not direct native
observations. The machine-readable
`graph_degree_source=requested_hnsw_same_graph_footprint_derivation` field makes
that provenance explicit, including when the writer uses its CPU fallback.

Both cuVS-backed algorithms can control sequential ingestion partitions and
the final segment topology. `premerge_segment_count`,
`force_merge_segment_count`, and `ram_per_thread_hard_limit_mb` must be
specified together. The partition count must be a positive integer, and
`force_merge_segment_count` must be either `0` (retain the partition segments)
or `1` (produce one final segment). CAGRA builds require
`force_merge_segment_count: 0`; they publish the segments built directly
rather than invoking the codec's vector-merge path.

Each equal, contiguous partition uses a separate writer lifecycle. Automatic
merges are disabled during ingestion, and the document-count flush boundary is
derived from the partition size. Values from 1 through 2047 MiB use Lucene's
public per-thread RAM-limit setter. If that limit causes an earlier flush, the
backend rejects the physical topology mismatch instead of silently reporting
the requested segment shape. Use smaller partitions when a partition cannot
fit under the public safety limit.

Larger direct-built segments are an explicit unsupported mode. A
`ram_per_thread_hard_limit_mb` value of 2048 MiB or greater also requires
`allow_unsupported_lucene_ram_limit: true` and
`force_merge_segment_count: 0`. The thin JAR applies the requested value to
Lucene 10.2's protected, non-public
`LiveIndexWriterConfig.perThreadHardLimitMB` field and reads the value back.
The bridge rejects a missing field, a field with the wrong type, an access or
write failure, or a getter-readback mismatch. The override only changes
Lucene's flush threshold: it does not reserve memory, make construction
out-of-core, guarantee that the host/JVM/GPU can hold the segment, or make a
later force merge safe.

When a final count of one is requested from more than one pre-merge segment, a
serial `forceMerge(1)` exercises the codec's vector-merge path; an index already
containing one segment does not invoke a merge and records
`runtime_force_merge_seconds=0` and `final_merge_policy=null`. This control does
not by itself make that merge out-of-core. For example, a four-partition build
that retains all four segments uses:

```yaml
name: lucene_accelerated_hnsw
groups:
m16_bw80_four_partitions:
build:
codec: ["Lucene101AcceleratedHNSWCodec"]
m: [16]
beam_width: [80]
premerge_segment_count: [4]
force_merge_segment_count: [0]
ram_per_thread_hard_limit_mb: [1945]
search: {}
```

This produces and retains four equal segments. Set
`force_merge_segment_count: [1]` to merge them serially into one segment after
ingestion. Results record the requested and observed pre-merge segment counts,
exact segment vector counts, derived `max_buffered_docs`, applied hard limit,
RAM-limit application mode, merge policies, and final segment count. Both RAM
limit paths verify the applied value at runtime. The manifest stores the
canonical request and runtime topology evidence. Reuse validates the evidence
against the request, dataset row count, and final physical segment count, and
reused build/search results surface that same evidence.

For example, this CAGRA configuration requests one directly built segment with
a 6144 MiB flush threshold:

```yaml
name: lucene_cuvs_cagra
groups:
one_large_segment:
build:
codec: ["CuVS2510GPUSearchCodec"]
premerge_segment_count: [1]
force_merge_segment_count: [0]
ram_per_thread_hard_limit_mb: [6144]
allow_unsupported_lucene_ram_limit: [true]
search: {}
```

Use this opt-in only after sizing the process and device for the complete
segment. The same topology keys are available to
`lucene_accelerated_hnsw`.

```bash
python -m cuvs_bench.run \
--backend lucene \
Expand All @@ -40,6 +147,12 @@ python -m cuvs_bench.run \

## Runtime requirements

Current indexes use manifest schema 4. Rebuild older manifests with
`--build --force`. Accelerated-HNSW index names now include the canonical `m`
and `beam_width` values, even when defaults are used. Old indexes are not
automatically migrated, and differently named indexes are not automatically
deleted or replaced.

The Lucene backend is opt-in because its runtime is not provisioned by the
ordinary cuVS Bench installation. Provisioning the required custom PyLucene
build is currently external to cuVS Bench. Every algorithm requires PyLucene
Expand All @@ -61,12 +174,38 @@ classpath inconsistent.
The initial backend accepts nonempty, finite, `float32` Euclidean/L2 vectors
and supports latency-mode sweeps. The CPU HNSW algorithm accepts at most 1024
dimensions; both cuVS-backed algorithms accept at most 4096. CAGRA uses the
codec's fixed defaults and supports `k <= 1024`. The backend validates the
codec's fixed defaults and supports `k <= 1024`. Accelerated HNSW accepts the
build parameters described above. The backend validates the
physical segment codec and every persisted vector field before searching, and
fails if CAGRA construction silently produced a brute-force index. Both HNSW
algorithms support an explicit `num_candidates` value greater than or equal to
`k`; their CPU search path is not subject to CAGRA's `k <= 1024` limit.

### Large-build memory and ingestion

The initial build path materializes the complete training-vector file as a
NumPy array. It then indexes one document at a time through PyLucene: every row
is converted to a Python list and Java `float[]`, a Lucene `Document` is
created, and `IndexWriter.addDocument` crosses the JCC boundary. This path is
not a streaming, bulk-FBIN, or out-of-core ingestion path.

Size the Python process and JVM heap for the dataset and codec being tested.
Additional JVM arguments can be supplied through the Lucene backend
configuration, for example:

```yaml
backend: lucene
jvm_args:
- -Xms16g
- -Xmx64g
- -XX:+ExitOnOutOfMemoryError
```

Pass the file with `--backend-config`. These values are illustrative, not
defaults: choose them from the available memory and expected workload. JVM
arguments are immutable after PyLucene initializes the process-global JVM, so
use a new Python process when changing them.

### Build PyLucene 10.2.0 from source

The ordinary cuVS Bench wheel, conda package, and container do not include the
Expand Down Expand Up @@ -173,6 +312,24 @@ JVM arguments cannot be changed after `lucene.initVM(...)`.

## Timing contract

`build_time_seconds` and `index_build_call_seconds` cover the in-process Lucene
build call: directory open/close, writer setup, document ingestion, an optional
synchronous force merge, writer commit/close, and the post-build reader check.
They exclude dataset loading, index verification, manifest publication,
installation, and size measurement; use
`backend_build_total_seconds` for that complete backend lifecycle.

In particular, `runtime_document_ingest_seconds` includes NumPy-to-Python and
Python-to-Java conversion, Java object creation, JCC dispatch, and Lucene work
performed by `addDocument`. It is not a measurement of cuvs-lucene or GPU graph
construction alone. `runtime_force_merge_seconds` measures an explicitly
requested synchronous `forceMerge` as a nested sub-timer; it is already included
in the enclosing build-call timings and must not be added to them.
`runtime_writer_commit_close_seconds` includes only work that the selected
codec defers until flush, commit, or close. Absolute build times from this
document-at-a-time path are therefore not directly comparable with a
Java-native or bulk-FBIN benchmark harness.

This initial backend invokes one Lucene query at a time. The common cuVS Bench
`--batch-size` value is retained for configuration and result-file
compatibility, but it does not introduce bulk or concurrent execution. Results
Expand Down Expand Up @@ -231,3 +388,29 @@ GPU; missing prerequisites fail rather than skip. Accelerated-HNSW GPU-intended
cases fail if the codec logs its CPU-writer fallback. A separate GPU-hidden
negative control verifies that this warning remains observable and attributable
to the case that produced it.

The direct PyLucene greater-than-2-GiB cases are independently gated because
they generate a 3 GiB FBIN and build one retained segment by issuing every
document through the Python/JCC boundary. Run them alone, in a fresh process,
on a local filesystem with at least 16 GiB free:

```bash
python -m pytest -q -s -x \
python/cuvs_bench/cuvs_bench/tests/test_lucene_large_segment_python_integration.py \
--run-lucene-large-segment-e2e \
--basetemp=/path/to/local-disk/lucene-large-pytest
```

Run this suite serially, without pytest-xdist. The `-x` option stops after the
first failure while pytest unwinds the fixture and removes its 3 GiB source
file.

The test configures a 12 GiB JVM maximum heap and requires at least 16 GiB of
currently available host memory. It maps the generated vectors read-only to
avoid a second 3 GiB Python allocation, but still exercises the production
per-document PyLucene conversion and `IndexWriter.addDocument` loop. Both the
CAGRA-built HNSW and GPU CAGRA cases require one physical segment and verify
the explicit 6144 MiB override. The HNSW case rejects logged CPU-writer
fallback; both cases reject brute-force-index fallback and graph-parameter
clamp warnings. The large-suite flag does not select the ordinary
live module, and `--run-lucene-e2e` does not select these capacity cases.
2 changes: 2 additions & 0 deletions fern/pages/lucene_api/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -14,6 +14,8 @@ For an introduction to the codecs, configuration, and tuning, see the [Lucene In
- [GPUIndex](/api-reference/lucene-api-com-nvidia-cuvs-lucene-gpuindex)
- [GPUSearchParams](/api-reference/lucene-api-com-nvidia-cuvs-lucene-gpusearchparams)
- [IndexSearcherTimingBridge](/api-reference/lucene-api-com-nvidia-cuvs-lucene-indexsearchertimingbridge)
- [IndexWriterConfigRAMLimitBridge](/api-reference/lucene-api-com-nvidia-cuvs-lucene-indexwriterconfigramlimitbridge)
- [Lucene101AcceleratedHNSWCodecFactory](/api-reference/lucene-api-com-nvidia-cuvs-lucene-lucene101acceleratedhnswcodecfactory)
- [LuceneProvider](/api-reference/lucene-api-com-nvidia-cuvs-lucene-luceneprovider)
- [ThreadLocalCuVSResourcesProvider](/api-reference/lucene-api-com-nvidia-cuvs-lucene-threadlocalcuvsresourcesprovider)
- [Utils](/api-reference/lucene-api-com-nvidia-cuvs-lucene-utils)
Expand Down
Original file line number Diff line number Diff line change
@@ -0,0 +1,32 @@
---
slug: api-reference/lucene-api-com-nvidia-cuvs-lucene-indexwriterconfigramlimitbridge
---

# IndexWriterConfigRAMLimitBridge

_Java package: `com.nvidia.cuvs.lucene`_

```java
public final class IndexWriterConfigRAMLimitBridge implements Function<Map<String, Object>, Map<String, Object>>
```

Applies and verifies Lucene's per-thread indexing-memory limit.

Lucene 10.2 accepts limits below 2048 MiB through its public setter. Larger limits require an
explicit opt-in and are applied to Lucene's non-public `perThreadHardLimitMB` field. That
unsupported override is intended only for controlled cuVS Bench vector-only builds that disable
automatic merges and validate their final segment topology. It must not be treated as a general
replacement for Lucene's safety limit.

The standard `Function` and `Map` types provide a narrow bridge for generated Java
bindings that do not expose reflection. The request must contain `config` (an
`IndexWriterConfig`), `per_thread_hard_limit_mb` (a positive `Integer`), and
`allow_unsupported_lucene_ram_limit` (a `Boolean`). The response returns the same
config, the verified limit, and `application_mode`, which is either `public_setter`
or `unsupported_field_override`.

The non-public path deliberately depends on Lucene's field name and type. It fails if the
field cannot be found, made accessible, written, or read back, so callers never silently continue
with a different limit.

_Source: `java/cuvs-lucene/src/main/java/com/nvidia/cuvs/lucene/IndexWriterConfigRAMLimitBridge.java:35`_
Original file line number Diff line number Diff line change
@@ -0,0 +1,54 @@
---
slug: api-reference/lucene-api-com-nvidia-cuvs-lucene-lucene101acceleratedhnswcodecfactory
---

# Lucene101AcceleratedHNSWCodecFactory

_Java package: `com.nvidia.cuvs.lucene`_

```java
public final class Lucene101AcceleratedHNSWCodecFactory implements Function<Map<String, Object>, Map<String, Object>>
```

Constructs an accelerated HNSW codec from one self-contained parameter request.

The standard `Function` and `Map` types form a narrow bridge for generated Java
bindings that do not wrap parameterized constructors. The request has exactly two entries,
max_conn and beam_width. Both must be `Integer` values in the inclusive range 1 through
512. The response contains the configured `codec` and the verified applied values under the
same parameter keys. Invalid requests throw `IllegalArgumentException`; codec construction
or initialization failures throw `IllegalStateException`.

The factory is stateless. In particular, it does not use JVM system properties, so concurrent
callers cannot observe or overwrite one another's configuration.

## Public Members

### apply

```java
@Override public Map<String, Object> apply(Map<String, Object> request)
```

Constructs one codec from the complete request.

**Parameters**

| Name | Description |
| --- | --- |
| `request` | exactly the `max_conn` and `beam_width` integer entries |

**Returns**

the codec and its verified applied parameter values

**Throws**

| Type | Description |
| --- | --- |
| `IllegalArgumentException` | if the request is null, incomplete, has extra keys, contains non-integer values, or contains values outside the inclusive range 1 through 512 |
| `IllegalStateException` | if codec construction or vector-format initialization fails |

_Source: `java/cuvs-lucene/src/main/java/com/nvidia/cuvs/lucene/Lucene101AcceleratedHNSWCodecFactory.java:42`_

_Source: `java/cuvs-lucene/src/main/java/com/nvidia/cuvs/lucene/Lucene101AcceleratedHNSWCodecFactory.java:28`_
Loading
Loading