Skip to content

Fix scalar CAGRA graph byte encoding - #15

Draft
nvzm123 wants to merge 1 commit into
mainfrom
zackm_scalar_cagra_encoding
Draft

nvzm123 wants to merge 1 commit into
mainfrom
zackm_scalar_cagra_encoding

Conversation

@nvzm123

@nvzm123 nvzm123 commented Oct 1, 2026 •

Copy link
Copy Markdown
Owner

Summary

Standalone fresh-main port of the scalar CAGRA input-encoding fix prototyped in fork PR #3. This branch starts at NVIDIA/cuvs main 313f8b548e2857139e9080b60aae5f6da3c52500; it does not include the earlier PR NVIDIA#2476 branch history.

  • Initialize scalar per-dimension extrema correctly, including all-negative dimensions.
  • Emit nonnegative 7-bit bytes (0–127) for the unsigned cuVS BYTE matrix. Remove the ineffective signed-to-unsigned byte copy.
  • Add deterministic quantizer regressions and a live GPU CAGRA-build → CPU HNSW-search test over two segments and a force merge. The live test checks GPU-writer activity, rank-one self hits, unique results, and at least 8/10 recall against brute-force Euclidean neighbors.
  • Regenerate the affected Lucene API pages.

Why and scope

The old (byte) (value & 0xff) conversion preserved the same bits; negative scalar codes were then interpreted as large unsigned values by cuVS. For dimensions whose old extrema were correct, moving the code by +64 restores the intended Euclidean pairwise distances without unsigned wrap. All-negative dimensions additionally needed the maximum initialization fix.

The public quantizer output values intentionally change. The on-disk format name, version, and read path do not change, so existing indexes should remain readable, although newly built graphs may have different topology. Mixed-version readability was not separately tested. This covers CAGRA-built HNSW with CPU search, not CAGRA search. Non-Euclidean quality and performance are not validated here; no benchmark selector or YAML API changes.

Validation

On NVIDIA A10G with Java 22, Maven 3.9.16, and matching cuVS/cuvs-java 26.12:

mvn -q -f java/cuvs-lucene/pom.xml spotless:check
mvn -q -f java/cuvs-lucene/pom.xml -Dtest=TestScalarQuantization,TestScalarQuantizedCagraGraph -Dcuvs.lucene.tests.requireGpu=true test
mvn -q -f java/cuvs-lucene/pom.xml -Dcuvs.lucene.tests.requireGpu=true clean verify
python3 fern/scripts/generate_api_reference.py --quiet
git diff --check

Spotless and generated-page checks passed. Focused tests: 3 run, 0 failed/errored/skipped. Full Maven verify: 350 run, 0 failed/errored, 30 skips from other tests. The cuvs.lucene.tests.requireGpu=true flag makes this scalar GPU test fail if cuVS is unavailable; without it, ordinary unsupported Maven runs skip this case. It does not change the prerequisite policy of other tests.

The focused native run emitted three initial-dimension (128 vs 0) warnings; the same warning class occurs in an unchanged main-branch CAGRA test. The full suite retains existing native/runtime/Maven warnings. Fern site validation was not run locally because Node.js is unavailable.

Pending larger-instance validation

Please rerun the property-bearing GPU test and representative Euclidean scalar quality/performance benchmarks on the larger EC2 instance before proposing this against NVIDIA/cuvs. Historical single-run 1M observations in fork PR #3 used a different integration base and are not validation of this head.

Integrated-stack benchmark results (2026-10-01)

These measurements used this scalar-encoding change applied as local commit 3dc0322fa on top of PR NVIDIA#2476 and PR NVIDIA#2653. They did not benchmark this standalone PR head (3f1d7228b) directly. Both codecs in the comparison used the same integrated stack.

Jasper-10M, 1536d, Euclidean; one segment, no force merge; graph degree 32 / intermediate degree 48; default IVF-PQ heuristics; efSearch=1500, topK=1500. Each index was searched twice in a fresh JVM (-Xms128m -Xmx512m), with 210 warmup and 1,000 measured queries per search. All index files were evicted before each repeat; the ANN vector and graph files were then prewarmed, while the scalar index's original float .vec was left cold. Indiscriminate index prewarming was disabled.

Codec Peak search-process RSS, two repeats Mean search latency, two repeats Recall@1500
Float CAGRA_HNSW 58.605 / 58.637 GiB 12.933 / 12.620 ms 98.0034%
Scalar CAGRA_HNSW_SCALAR 15.694 / 15.691 GiB 8.326 / 8.312 ms 97.8883%

Scalar reduced the hot search-process RSS by 42.93 GiB (73.2%) on average across repeats, while both indexes exceeded 95% recall. The difference was predominantly file-backed mmap data, not Java heap: fincore showed 57.220 GiB of float .vec resident versus 14.342 GiB of scalar .veq, with only 8 KiB of scalar .vec resident. These pages are OS-reclaimable, but compete for DRAM when hot. The raw float .vec remains on disk; blanket prewarming or exact-vector access can make it resident. The latency values are repeats on one built index per codec, not a confidence interval.

For the matched cold-source builds (-Xms64g -Xmx256g), scalar peak build-phase RSS was 96.841 GiB versus 124.927 GiB for float; indexing took 332.236 s versus 268.413 s, and index size was 72.550 GiB versus 58.207 GiB. These single-build timings should not be overinterpreted.

To isolate this PR's removal of the redundant byte-array clone within the scalar writer, a local copy-only ablation retained the corrected quantizer but restored the old, bit-preserving clone. Two cold Jasper scalar builds per variant measured peak build RSS of 96.976 / 96.893 GiB without the clone versus 107.007 / 107.381 GiB with it: a median 10.26 GiB (9.6%) saving. This PR does not change the search reader, so the 42.93 GiB scalar-versus-float search saving is a codec-selection result, not an incremental search-memory gain caused by this PR alone. These Jasper results do not establish quality-equivalent scalar behavior for Deep1B-100M, which did not reach 95% recall in earlier tests at these settings.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant