From fec567f804b80875b5f41c03e89e109d13791a56 Mon Sep 17 00:00:00 2001 From: nvzm123 Date: Tue, 6 Oct 2026 04:31:07 +0000 Subject: [PATCH 1/2] Document Lucene benchmark backend --- fern/docs.yml | 5 +- fern/pages/cuvs_bench/lucene_backend.md | 238 ++++++++++++++++++++++++ fern/pages/cuvs_bench/running.md | 1 + 3 files changed, 243 insertions(+), 1 deletion(-) create mode 100644 fern/pages/cuvs_bench/lucene_backend.md diff --git a/fern/docs.yml b/fern/docs.yml index 921d86ce70..c5b3fd6842 100644 --- a/fern/docs.yml +++ b/fern/docs.yml @@ -219,8 +219,11 @@ navigation: path: "./pages/cuvs_bench/datasets.md" - page: "Synthetic Dataset Generation" path: "./pages/cuvs_bench/synthesize_dataset.md" - - page: "Backends" + - section: "Backends" path: "./pages/cuvs_bench/pluggable_backend.md" + contents: + - page: "Lucene" + path: "./pages/cuvs_bench/lucene_backend.md" - page: "cuVS Bench Parameter Tuning Guide" hidden: true path: "./pages/cuvs_bench/param_tuning.md" diff --git a/fern/pages/cuvs_bench/lucene_backend.md b/fern/pages/cuvs_bench/lucene_backend.md new file mode 100644 index 0000000000..1c6782632d --- /dev/null +++ b/fern/pages/cuvs_bench/lucene_backend.md @@ -0,0 +1,238 @@ +--- +slug: user-guide/benchmarking-guide/cu-vs-bench-tool/lucene-backend +--- + +# Lucene Backend + +The optional `lucene` backend runs cuVS Bench against a local Lucene index. It +embeds the JVM through PyLucene. The CPU HNSW control uses stock Lucene through +PyLucene. The cuVS-Lucene thin JAR supplies GPU CAGRA search and GPU-accelerated +HNSW construction with CPU HNSW search. PyLucene is an implementation detail; +the public cuVS Bench backend name is `lucene`. + +Selecting this backend is explicit. Commands that do not select it through a +Lucene `--backend-config` file continue to use the `cpp_gbench` backend and its +existing `cuvs_cagra` default. Save this reusable selector as +`lucene-backend.yaml`: + +```yaml +backend: lucene +``` + +## Algorithms + +| Algorithm | Codec | Build and search path | +| --- | --- | --- | +| `lucene_cuvs_cagra` | `CuVS2510GPUSearchCodec` | cuVS CAGRA on a supported NVIDIA GPU | +| `lucene_accelerated_hnsw` | `Lucene101AcceleratedHNSWCodec` | cuVS CAGRA build with Lucene CPU HNSW search; CPU HNSW build fallback when GPU support is unavailable | +| `lucene_cpu_hnsw` | `Lucene101` | Lucene CPU HNSW control | + +`lucene_cuvs_cagra` is selected when the Lucene backend configuration is used +without an explicit `--algorithms` value. Select `lucene_cpu_hnsw` explicitly +when a CPU control is needed. + +```bash +python -m cuvs_bench.run \ + --backend-config lucene-backend.yaml \ + --dataset test-data \ + --dataset-path /absolute/path/to/datasets \ + --algorithms lucene_cuvs_cagra \ + --groups test \ + --batch-size 10 \ + -k 10 \ + --build --search +``` + +## Runtime requirements + +The Lucene backend is opt-in because its runtime is not provisioned by the +ordinary cuVS Bench installation. Provisioning the required custom PyLucene +build is currently external to cuVS Bench. Every algorithm requires PyLucene +10.2.0 and JDK 22. Both cuVS-backed algorithms require matching `cuvs-java` +and thin `cuvs-lucene` JARs. GPU execution additionally requires compatible +native cuVS, CUDA, and a supported NVIDIA GPU. `lucene_cuvs_cagra` fails when +that GPU path is unavailable. `lucene_accelerated_hnsw` instead retains the +codec's intentional CPU-writer fallback. +The backend discovers artifacts built by the current checkout or installed in +the local Maven repository. For nonstandard locations, set both +`CUVS_LUCENE_CUVS_JAVA_JAR` and `CUVS_LUCENE_JAR`; use `JAVA_LIBRARY_PATH` for +native libraries. An explicit native-library path must contain an unversioned +`libcuvs_c.so`. + +Do not use the cuVS-Lucene JAR assembled with dependencies: PyLucene already +provides Lucene classes, and loading a second copy can make the embedded JVM +classpath inconsistent. + +The initial backend accepts nonempty, finite, `float32` Euclidean/L2 vectors +and supports latency-mode sweeps. The CPU HNSW algorithm accepts at most 1024 +dimensions; both cuVS-backed algorithms accept at most 4096. CAGRA uses the +codec's fixed defaults and supports `k <= 1024`. The backend validates the +physical segment codec and every persisted vector field before searching, and +fails if CAGRA construction silently produced a brute-force index. Both HNSW +algorithms support an explicit `num_candidates` value greater than or equal to +`k`; their CPU search path is not subject to CAGRA's `k <= 1024` limit. + +### Build PyLucene 10.2.0 from source + +The ordinary cuVS Bench wheel, conda package, and container do not include the +custom PyLucene 10.2.0 runtime. The helper below is available only in a cuVS +source checkout. Automated CI provisioning for the live integration suite is +tracked in [NVIDIA/cuvs#2635](https://github.com/NVIDIA/cuvs/issues/2635). + +This procedure has been validated on Linux x86_64 with CPython 3.14 and JDK 22. +A PyLucene wheel is specific to its operating system, architecture, Python ABI, +and JDK toolchain; do not attach or redistribute this local wheel as a general +binary. No real ARM64 build or execution has been performed. + +Start in the normal cuVS source-build environment described in the +[shared source-build prerequisites](/installation#build-from-source). Also +install the [Java build prerequisites](/installation/java#build-from-source), +CPython 3.11-3.14 with development headers and `venv` support, GNU Make, a C/C++ +compiler, `awk`, `curl`, `flock`, `gzip`, `patch`, `tar`, GNU coreutils, and GNU +findutils. CUDA and a supported NVIDIA GPU are required for the GPU integration +cases. + +To run either cuVS-backed algorithm or the full CPU/GPU integration suite, start +at the repository root and, before activating the isolated PyLucene environment, +build matching native cuVS, base `cuvs-java`, and thin `cuvs-lucene` artifacts. +The artifact-free `lucene_cpu_hnsw` control does not need native cuVS or either +JAR, so CPU-only users can skip this command and the native-library setup below, +but should retain the documented cuVS build environment for the editable install. + +```bash +export JAVA_HOME=/absolute/path/to/jdk-22 +./build.sh libcuvs java lucene +``` + +Choose a stable absolute PyLucene build location outside `/tmp`. Do not move +the completed directory or selected JDK, and do not remove the base Python +installation: JCC and the virtual environment retain absolute paths. Allow at +least 3 GB of disk space. + +```bash +export PYLUCENE_BUILD_ROOT="$HOME/.local/share/cuvs/pylucene-10.2.0" + +conda/recipes/cuvs-bench/build_pylucene_10_2.sh \ + --python python3 \ + --build-root "$PYLUCENE_BUILD_ROOT" + +source "$PYLUCENE_BUILD_ROOT/activate.sh" +python -m pip install -e ./python/cuvs_bench +python -m pip check +``` + +The helper verifies checksum-pinned Apache PyLucene 10.0.0 scaffolding, Lucene +10.2.0 sources, and the Gradle distribution. It applies the tracked +compatibility patch, builds JCC 3.15 and a Python-ABI-specific PyLucene wheel in +an isolated virtual environment, runs the upstream PyLucene tests, and performs +a JVM class-loading smoke test. The Python packages requested by the helper are +version-pinned but not hash-locked, and Gradle dependencies remain +network-resolved, so this is not a hermetic or bit-for-bit-reproducible build. + +Verify that the activated interpreter uses the expected runtime: + +```bash +python - <<'PY' +import os +from pathlib import Path + +import lucene + +build_root = Path(os.environ["PYLUCENE_BUILD_ROOT"]).resolve() +module_path = Path(lucene.__file__).resolve() +assert lucene.VERSION == "10.2.0", lucene.VERSION +assert module_path.is_relative_to(build_root), module_path +print(f"PyLucene {lucene.VERSION} from {module_path}") +PY +``` + +The cuVS-backed algorithms do not require the `bench-ann` target. Make the +fresh source-build libraries visible to both Java and the ELF loader: + +```bash +CUVS_NATIVE_BUILD="$(cd cpp/build && pwd -P)" +export JAVA_LIBRARY_PATH="$CUVS_NATIVE_BUILD/c:$CUVS_NATIVE_BUILD" +export LD_LIBRARY_PATH="$JAVA_LIBRARY_PATH${LD_LIBRARY_PATH:+:$LD_LIBRARY_PATH}" +``` + +Keep CUDA's runtime directory loader-visible through the normal cuVS build +environment. Some conda toolchains encode their active prefix as `DT_RPATH`, +which takes precedence over `LD_LIBRARY_PATH`. The standard `./build.sh` +command installs the freshly built libraries into that active prefix. If you +override `INSTALL_PREFIX`, verify that the C wrapper resolves matching cuVS +libraries and relocations rather than an older installation: + +```bash +ldd -r "$CUVS_NATIVE_BUILD/c/libcuvs_c.so" \ + | grep -E 'libcuvs|librmm|librapids_logger|undefined symbol' +``` + +The backend discovers the matching JARs under the current checkout; do not +select the native-classifier `cuvs-java` JAR or the cuVS-Lucene +`jar-with-dependencies`. + +In each new shell, reactivate the normal cuVS source-build environment before +sourcing `activate.sh`, then restore the two native-library path variables +before starting Python. PyLucene's JVM is process-global, so its classpath and +JVM arguments cannot be changed after `lucene.initVM(...)`. + +## Timing contract + +This initial backend invokes one Lucene query at a time. The common cuVS Bench +`--batch-size` value is retained for configuration and result-file +compatibility, but it does not introduce bulk or concurrent execution. Results +therefore record both `requested_batch_size` and +`effective_search_batch_size=1` and use one latency sample per query. + +The headline latency is the complete `client_query` boundary: Java vector and +query preparation, the synchronous search call, and hit materialization. QPS +uses the wall time for the entire serial query corpus, so it is not simply the +reciprocal of mean latency. The raw result also reports narrower preparation, +PyLucene search-dispatch, result-materialization, reader lifecycle, and NumPy +conversion timings. `search_dispatch_kind` distinguishes direct PyLucene +dispatch from the thin-JAR timing bridge. Compare timing results only when that +value matches. Selecting the CPU and GPU algorithms together loads the bridge +for both and provides a like-for-like dispatch-boundary comparison; the +artifact-free CPU control remains useful but is not dispatch-identical. + +When the backend initializes from a validated cuvs-java and cuvs-lucene +artifact pair, an additional `java_index_searcher_search` measurement uses +`System.nanoTime()` around exactly `IndexSearcher.search(Query, int)` inside +the JVM. This exact Java measurement is diagnostic and is unavailable to the +artifact-free CPU control. The reported timing boundaries are nested rather +than additive; summing parent and child values double-counts work. + +`backend_search_invocation_total_ms` measures one complete non-dry-run search +invocation after argument validation and dry-run handling. A parameter sweep +copies that one shared invocation value into every plan row; do not interpret +or sum it as a per-plan duration. + +Before each search-parameter plan, the backend sequentially reads every regular +index file to establish the same host page-cache policy for CPU and GPU paths. +It runs no discarded warmup queries: the first query remains part of the +result, and first-query and subsequent-query values are reported separately +for the client-query and PyLucene-dispatch boundaries, plus the exact JVM +search boundary when Java timing is available. That contrast can reveal +first-query effects, but it does not establish steady state or isolate every +JVM, CUDA, or cuVS initialization cost. Build results similarly separate +dataset preparation, runtime setup, writer lifecycle, index validation, +publication, and total backend time. + +## Integration tests + +Run the live CPU/GPU suite from the repository root after providing the same +runtime prerequisites: + +```bash +python -m pytest -q -s \ + python/cuvs_bench/cuvs_bench/tests/test_lucene_integration.py \ + --run-lucene-e2e +``` + +Without `--run-lucene-e2e`, ordinary cuVS Bench test runs skip these live +cases. Once selected, every case requires PyLucene and Java. GPU-intended cases +also require the cuVS Java artifacts, native libraries, CUDA, and a supported +GPU; missing prerequisites fail rather than skip. Accelerated-HNSW GPU-intended +cases fail if the codec logs its CPU-writer fallback. A separate GPU-hidden +negative control verifies that this warning remains observable and attributable +to the case that produced it. diff --git a/fern/pages/cuvs_bench/running.md b/fern/pages/cuvs_bench/running.md index 94606bb214..cf1efd1866 100644 --- a/fern/pages/cuvs_bench/running.md +++ b/fern/pages/cuvs_bench/running.md @@ -63,6 +63,7 @@ Create a custom YAML file with a `base` group to override the default benchmark | GGNN | `ggnn` | | HNSWLIB | `hnswlib` | | DiskANN | `diskann_memory`, `diskann_ssd` | +| Lucene (`backend: lucene` config) | `lucene_cpu_hnsw`, `lucene_accelerated_hnsw`, `lucene_cuvs_cagra` | | NVIDIA cuVS | `cuvs_brute_force`, `cuvs_cagra`, `cuvs_ivf_flat`, `cuvs_ivf_pq`, `cuvs_cagra_hnswlib`, `cuvs_vamana` | ### Multi-GPU algorithms From d5337aa3df66c6e0d0a8106883e97ce3247f61e9 Mon Sep 17 00:00:00 2001 From: nvzm123 Date: Tue, 6 Oct 2026 21:15:56 +0000 Subject: [PATCH 2/2] Document pluggable backend contracts --- fern/pages/cuvs_bench/pluggable_backend.md | 251 ++++++++++++--------- 1 file changed, 149 insertions(+), 102 deletions(-) diff --git a/fern/pages/cuvs_bench/pluggable_backend.md b/fern/pages/cuvs_bench/pluggable_backend.md index ce5eb2ec51..5b49b45385 100644 --- a/fern/pages/cuvs_bench/pluggable_backend.md +++ b/fern/pages/cuvs_bench/pluggable_backend.md @@ -31,7 +31,8 @@ results = orchestrator.run_benchmark( 2. The orchestrator finds the config loader registered for the requested backend type. 3. The config loader returns a `DatasetConfig` and one or more `BenchmarkConfig` objects. 4. The orchestrator creates the backend registered for the same backend type. -5. The backend runs `build(...)` and `search(...)`, then returns `BuildResult` and `SearchResult` objects. +5. The backend runs `build(...)` and `search(...)`, then returns a `BuildResult` + and a list of `SearchResult` objects. The config loader decides what to run. The backend decides how to run it. @@ -42,7 +43,7 @@ A config loader receives the arguments passed to `run_benchmark()`, such as `dat | Return value | Purpose | | --- | --- | | `DatasetConfig` | Dataset metadata, including vector files, ground-truth files, distance metric, dimensions, and optional subset size. | -| `List[BenchmarkConfig]` | One or more benchmark configurations. Each contains index configurations and backend-specific options. | +| `list[BenchmarkConfig]` | One or more benchmark configurations. Each contains index configurations and backend-specific options. | Each `IndexConfig` describes one index to benchmark: @@ -54,109 +55,112 @@ Each `IndexConfig` describes one index to benchmark: | `search_params` | Search parameter combinations to benchmark. | | `file` | Path or identifier where the backend stores the index. | -The following minimal loader creates one dataset, one index, and one search configuration: +`ConfigLoader.load()` owns dataset loading and parameter expansion. A backend +loader implements the two template hooks below. This minimal network loader +creates one index configuration for each requested algorithm and group: ```python +from pathlib import Path + from cuvs_bench.orchestrator.config_loaders import ( - ConfigLoader, - DatasetConfig, BenchmarkConfig, + ConfigLoader, IndexConfig, ) class MyConfigLoader(ConfigLoader): + def __init__(self, config_path=None): + self.config_path = str( + Path(config_path) if config_path else Path(__file__).parent / "config" + ) + @property def backend_type(self) -> str: return "my_backend" - def load(self, dataset, dataset_path, algorithms, count=10, batch_size=10000, **kwargs): - dataset_config = DatasetConfig( - name=dataset, - base_file=..., - query_file=..., - groundtruth_neighbors_file=..., - distance="euclidean", - dims=128, - ) - index = IndexConfig( - name=f"{algorithms}.default", - algo=algorithms, - build_param={"nlist": 1024}, - search_params=[{"nprobe": 10}], - file=..., - ) - benchmark_config = BenchmarkConfig( - indexes=[index], - backend_config={ - "host": ..., - "port": ..., - "index_name": ..., - }, - ) - return dataset_config, [benchmark_config] + def _discover_algo_groups( + self, dataset_conf, dataset, dataset_path, **kwargs + ): + algorithms = (kwargs.get("algorithms") or "my_algorithm").split(",") + groups = (kwargs.get("groups") or "default").split(",") + group_config = { + "build": {}, + "search": {"ef_search": [100]}, + } + return [ + (algorithm.strip(), group.strip(), group_config, {}) + for algorithm in algorithms + for group in groups + ] + + def _build_benchmark_configs( + self, + dataset_config, + dataset_conf, + dataset, + dataset_path, + expanded_groups, + **kwargs, + ): + configurations = [] + for ( + algorithm, + group, + _group_config, + build_combinations, + search_combinations, + _group_metadata, + ) in expanded_groups: + indexes = [ + IndexConfig( + name=f"{algorithm}.{group}", + algo=algorithm, + build_param=build_parameters, + search_params=search_combinations, + file=str( + Path(dataset_path) + / dataset + / "index" + / f"{algorithm}.{group}" + ), + ) + for build_parameters in build_combinations + ] + configurations.append( + BenchmarkConfig( + indexes=indexes, + backend_config={ + "host": kwargs.get("host", "localhost"), + "port": kwargs.get("port", 9200), + "algo": algorithm, + }, + ) + ) + return configurations ``` ## Adding a backend Add a new backend when cuVS Bench needs to drive a different execution path, such as a vector database, remote service, or custom benchmark runner. -1. Implement a config loader by subclassing `ConfigLoader` from `cuvs_bench.orchestrator.config_loaders`. Its `load()` method should return `(DatasetConfig, List[BenchmarkConfig])`. -2. Implement a backend by subclassing `BenchmarkBackend` from `cuvs_bench.backends.base`. Its `build()` method should return `BuildResult`; its `search()` method should return `SearchResult`. +1. Implement a config loader by subclassing `ConfigLoader` and defining + `_discover_algo_groups()`, `_build_benchmark_configs()`, and `backend_type`. + The inherited `load()` method returns + `(DatasetConfig, list[BenchmarkConfig])`. +2. Implement a backend by subclassing `BenchmarkBackend`. Its `build()` method + returns `BuildResult`; its `search()` method returns `list[SearchResult]`. 3. Register both pieces with the same backend type name. -4. Run benchmarks with `BenchmarkOrchestrator(backend_type="my_backend")`. +4. Select an installed backend with a YAML `backend:` field passed through + `--backend-config`, or construct `BenchmarkOrchestrator` directly with its + `backend_type`. -```python -from cuvs_bench.orchestrator import register_config_loader -from cuvs_bench.backends import get_registry - -register_config_loader("my_backend", MyConfigLoader) -get_registry().register("my_backend", MyBackend) -``` +The complete example below uses an idempotent function to register both pieces. -## Example: Elasticsearch backend +## Example: network backend -This example shows the shape of a network backend. The loader creates the dataset and benchmark configs. The backend uses `backend_config` to connect to the service, build the index, run search, and return cuVS Bench result objects. - -```python -from cuvs_bench.orchestrator.config_loaders import ( - ConfigLoader, - DatasetConfig, - BenchmarkConfig, - IndexConfig, -) - -class ElasticsearchConfigLoader(ConfigLoader): - @property - def backend_type(self) -> str: - return "elasticsearch" - - def load(self, dataset, dataset_path, algorithms, count=10, batch_size=10000, **kwargs): - dataset_config = DatasetConfig( - name=dataset, - base_file=..., - query_file=..., - groundtruth_neighbors_file=..., - distance="euclidean", - dims=kwargs.get("dims", 128), - ) - index = IndexConfig( - name=f"{algorithms}.es", - algo=algorithms, - build_param={}, - search_params=[{"ef_search": 100}], - file=..., - ) - benchmark_config = BenchmarkConfig( - indexes=[index], - backend_config={ - "host": ..., - "port": ..., - "index_name": ..., - "algo": algorithms, - }, - ) - return dataset_config, [benchmark_config] -``` +This example shows the shape of a network backend using the loader above. The +backend consumes `backend_config` to connect to the service, build the index, +run search, and return cuVS Bench result objects. ```python import numpy as np @@ -166,10 +170,12 @@ from cuvs_bench.backends.base import ( SearchResult, ) -class ElasticsearchBackend(BenchmarkBackend): +class MyBackend(BenchmarkBackend): + default_algorithm = "my_algorithm" + @property def algo(self) -> str: - return self.config.get("algo", "elasticsearch") + return self.config.get("algo", "my_algorithm") def build(self, dataset, indexes, force=False, dry_run=False): return BuildResult( @@ -194,33 +200,74 @@ class ElasticsearchBackend(BenchmarkBackend): dry_run=False, ): n_queries = dataset.n_queries - return SearchResult( - neighbors=np.zeros((n_queries, k), dtype=np.int64), - distances=np.zeros((n_queries, k), dtype=np.float32), - search_time_ms=0.0, - queries_per_second=0.0, - recall=0.0, - algorithm=self.algo, - search_params=indexes[0].search_params if indexes else [], - success=True, - ) + return [ + SearchResult( + neighbors=np.zeros((n_queries, k), dtype=np.int64), + distances=np.zeros((n_queries, k), dtype=np.float32), + search_time_ms=0.0, + queries_per_second=0.0, + recall=0.0, + algorithm=self.algo, + search_params=indexes[0].search_params if indexes else [], + success=True, + ) + ] ``` ```python -from cuvs_bench.orchestrator import register_config_loader -from cuvs_bench.backends import get_registry +from cuvs_bench.backends.registry import ( + get_registry, + list_config_loaders, + register_backend, + register_config_loader, +) -register_config_loader("elasticsearch", ElasticsearchConfigLoader) -get_registry().register("elasticsearch", ElasticsearchBackend) +def register(): + registry = get_registry() + if not registry.is_registered("my_backend"): + register_backend("my_backend", MyBackend) + if "my_backend" not in list_config_loaders(): + register_config_loader("my_backend", MyConfigLoader) ``` +Installed plugins publish the same idempotent registration function in both +entry-point groups: + +```toml +[project.entry-points."cuvs_bench.backends"] +my_backend = "my_package.backend:register" + +[project.entry-points."cuvs_bench.config_loaders"] +my_backend = "my_package.backend:register" +``` + +Select the installed backend with a small YAML file: + +```yaml +backend: my_backend +host: localhost +port: 9200 +``` + +```bash +python -m cuvs_bench.run --backend-config my-backend.yaml +``` + +`BenchmarkBackend.default_algorithm` supplies the CLI's algorithm default for +the selected backend. The default value is `None`, which retains the cuVS Bench +CLI default. A backend may also override `result_failure_message(results)` to +return a message when failed result objects should make the CLI exit nonzero +after export processing. Unsuccessful, skipped, and dry-run results are +excluded from the CSV output and do not replace prior measurements. Returning +`None` keeps the default nonfatal policy. + ## Components at a glance | Component | Description | | --- | --- | -| `ConfigLoader` | Abstract class whose `load(**kwargs)` method returns `(DatasetConfig, List[BenchmarkConfig])`. Register with `register_config_loader(backend_type, loader_class)`. | -| `BenchmarkBackend` | Abstract class whose `build(...)` method returns `BuildResult` and whose `search(...)` method returns `SearchResult`. Register with `BackendRegistry.register(name, backend_class)`. | -| `BackendRegistry` | Singleton registry returned by `get_registry()`. It maps backend type names to backend classes. | +| `ConfigLoader` | Template base class whose `load(**kwargs)` method returns `(DatasetConfig, list[BenchmarkConfig])`; plugins implement its discovery/build hooks. Register with `register_config_loader(backend_type, loader_class)`. | +| `BenchmarkBackend` | Abstract class whose `build(...)` method returns `BuildResult` and whose `search(...)` method returns `list[SearchResult]`. Its optional class hooks control the CLI algorithm default and fatal-result policy. | +| `BackendRegistry` | Singleton registry returned by `get_registry()` that stores registered backend classes. The `get_backend_class()` and `get_config_loader()` helpers discover matching installed entry points on demand. | ## C++ Backend