Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
19 commits
Select commit Hold shift + click to select a range
87a0075
Migrate CAGRA merge() to caller-concatenated dataset + offsets contract
HowardHuang1 Aug 26, 2026
50dd606
Merge remote-tracking branch 'upstream/main' into HH-migrate-merge-da…
HowardHuang1 Aug 26, 2026
3210112
Merge remote-tracking branch 'upstream/main' into HH-migrate-merge-da…
HowardHuang1 Aug 27, 2026
0996cb1
Fix Rust/Go CAGRA merge CI failures
HowardHuang1 Aug 27, 2026
a83394f
Merge remote-tracking branch 'upstream/main' into HH-migrate-merge-da…
HowardHuang1 Aug 27, 2026
1cede24
Merge remote-tracking branch 'upstream/release/26.10' into HH-migrate…
HowardHuang1 Sep 10, 2026
512e8b8
Merge remote-tracking branch 'upstream/release/26.10' into HH-migrate…
HowardHuang1 Sep 14, 2026
8054382
Merge remote-tracking branch 'upstream/release/26.10' into HH-migrate…
HowardHuang1 Sep 14, 2026
250b4da
remove kvikio dependency by using public build() function rather than…
HowardHuang1 Sep 14, 2026
7a4845d
Fix failing cuvs lucene tests
HowardHuang1 Sep 14, 2026
a14a1da
add fern changes from fern pre-commit hooks so that style-checker on …
HowardHuang1 Sep 15, 2026
ab5fe77
Merge remote-tracking branch 'upstream/release/26.10' into HH-migrate…
HowardHuang1 Sep 20, 2026
a8e86d7
Regenerate Fern Lucene writer API docs after merge line-number shift
HowardHuang1 Sep 21, 2026
cf27fa0
Merge remote-tracking branch 'upstream/release/26.10' into HH-migrate…
HowardHuang1 Sep 23, 2026
714c749
Add concatenate_datasets()/concatenate_and_filter_datasets() CAGRA me…
HowardHuang1 Sep 23, 2026
2206171
Add C API for CAGRA concatenate_datasets()/concatenate_and_filter_dat…
HowardHuang1 Sep 23, 2026
7df7f51
Add CAGRA concatenateDatasets()/concatenateAndFilterDatasets() to Jav…
HowardHuang1 Sep 29, 2026
3d8592b
Restore deprecated cuvsCagraMerge/cuvsCagraMergeWithParams for C ABI …
HowardHuang1 Oct 1, 2026
e5e16ad
Update Go/Python/Rust bindings to call cuvsCagraMerge_v2 / cuvsCagraM…
HowardHuang1 Oct 1, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
212 changes: 178 additions & 34 deletions c/include/cuvs/neighbors/cagra.h
Original file line number Diff line number Diff line change
Expand Up @@ -1012,9 +1012,89 @@ CUVS_EXPORT cuvsError_t cuvsCagraIndexFromArgs(cuvsResources_t res,
* @{
*/

/**
* @brief Compute per-index write offsets for a merged dataset buffer.
*
* `cuvsCagraMerge_v2`/`cuvsCagraMergeWithParams_v2` require the caller to have already
* concatenated every input index's dataset (in `indices` order, applying `filter` if any) into a
* single buffer and to know each index's starting row within it. For `filter.type == NO_FILTER`,
* those offsets are just the cumulative sizes of `indices` and this function is not needed. For
* `BITSET`, the number of surviving rows per index cannot be derived any other way, so call this
* first.
*
* @param[in] res cuvsResources_t opaque C handle
* @param[in] indices Array of input cuvsCagraIndex_t handles that will be passed to merge
* @param[in] num_indices Number of input indices
* @param[in] filter Filter that will be passed to merge. Only `NO_FILTER` and `BITSET` supported.
* @param[out] offsets Caller-allocated array of `num_indices + 1` int64_t. Entry `i` is the row at
* which `indices[i]`'s surviving rows must start in the merged buffer; the
* last entry is the total row count of the merged buffer.
* @return cuvsError_t
*/
CUVS_EXPORT cuvsError_t cuvsCagraMergedDatasetOffsets(cuvsResources_t res,
cuvsCagraIndex_t* indices,
size_t num_indices,
cuvsFilter filter,
int64_t* offsets);

/**
* @brief Concatenate every input index's dataset (unfiltered, in `indices` order) into a freshly
* allocated, CAGRA-padded, owning device dataset.
*
* This is an optional convenience function for building `cuvsCagraMerge_v2`'s `merged_dataset`
* argument in the unfiltered case; callers that already have their own concatenated buffer are not
* required to use it. The matching `offsets` are simply each index's cumulative row count, or
* `cuvsCagraMergedDatasetOffsets` called with `filter.type == NO_FILTER`.
*
* @param[in] res cuvsResources_t opaque C handle
* @param[in] indices Array of input cuvsCagraIndex_t handles to concatenate, in the order they
* will be passed to `cuvsCagraMerge_v2`
* @param[in] num_indices Number of input indices
* @param[out] merged_dataset Newly allocated owning dataset handle containing every index's rows,
* concatenated in `indices` order. Caller must destroy it with
* `cuvsDatasetDestroy` when done.
* @return cuvsError_t
*/
CUVS_EXPORT cuvsError_t cuvsCagraConcatenateDatasets(cuvsResources_t res,

@cjnolet cjnolet Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't see why this is CAGRA-specific. This looks like a dataset-specific function. Please move to dataset.h.

cuvsCagraIndex_t* indices,
size_t num_indices,
cuvsDataset_t* merged_dataset);

/**
* @brief Concatenate every input index's dataset (in `indices` order), retaining only the rows
* selected by `filter`, into a freshly allocated, CAGRA-padded, owning device dataset.
*
* This is an optional convenience function for building `cuvsCagraMerge_v2`'s `merged_dataset`
* argument in the bitset-filtered case; callers that already have their own filtered, concatenated
* buffer are not required to use it. Call `cuvsCagraMergedDatasetOffsets` with the same `filter`
* to get the matching `offsets`.
*
* @param[in] res cuvsResources_t opaque C handle
* @param[in] indices Array of input cuvsCagraIndex_t handles to concatenate, in the order they
* will be passed to `cuvsCagraMerge_v2`
* @param[in] num_indices Number of input indices
* @param[in] filter Bitset row filter selecting which rows survive into the output. Only `BITSET`
* is supported.
* @param[out] merged_dataset Newly allocated owning dataset handle containing every index's
* surviving rows, concatenated in `indices` order. Caller must destroy
* it with `cuvsDatasetDestroy` when done.
* @return cuvsError_t
*/
CUVS_EXPORT cuvsError_t cuvsCagraConcatenateAndFilterDatasets(cuvsResources_t res,

@cjnolet cjnolet Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's simplify this and just have 1 function for concatenating which accepts a filter (aka consolidate this and the function above).

cuvsCagraIndex_t* indices,
size_t num_indices,
cuvsFilter filter,
cuvsDataset_t* merged_dataset);

/**
* @brief Merge multiple CAGRA indices into a single CAGRA index.
*
* The caller is responsible for concatenating every input index's dataset (applying `filter` if
* any) into a single `merged_dataset` buffer before calling this, and for computing `offsets`
* (see `cuvsCagraMergedDatasetOffsets`). This function only builds/merges the graph and rebinds
* the output index to `merged_dataset` -- it never allocates or copies dataset rows itself. This
* mirrors the `cuvsCagraExtend` contract.
*
* All input indices must have been built with the same data type (`index.dtype`) and
* have the same dimensionality (`index.dims`). The merged index uses the output
* parameters specified in `cuvsCagraIndexParams`. The merge algorithm is selected automatically.
Expand Down Expand Up @@ -1047,12 +1127,14 @@ CUVS_EXPORT cuvsError_t cuvsCagraIndexFromArgs(cuvsResources_t res,
* cuvsCagraIndexParams_t merge_params;
* cuvsError_t params_create_status = cuvsCagraIndexParamsCreate(&merge_params);
*
* cuvsDataset_t merged_dataset;
* cuvsDatasetCreate(&merged_dataset);
* // Build `merged_dataset` as the caller-owned concatenation of index1 || index2 (e.g. via
* // cuvsDatasetMakePadded over a device buffer you populated yourself).
* cuvsDataset_t merged_dataset = ...;
* int64_t offsets[3] = {0, index1_size, index1_size + index2_size};
* cuvsFilter filter = {.type = NO_FILTER, .addr = 0};
*
* cuvsError_t merge_status = cuvsCagraMerge(res, merge_params, (cuvsCagraIndex_t[]){index1,
* index2}, 2, filter, merged_dataset, merged_index);
* cuvsError_t merge_status = cuvsCagraMerge_v2(res, merge_params, (cuvsCagraIndex_t[]){index1,
* index2}, 2, filter, merged_dataset, offsets, merged_index);
*
* // Use merged_index for search operations
*
Expand All @@ -1066,51 +1148,113 @@ CUVS_EXPORT cuvsError_t cuvsCagraIndexFromArgs(cuvsResources_t res,
* @param[in] params cuvsCagraIndexParams_t parameters for the output index
* @param[in] indices Array of input cuvsCagraIndex_t handles to merge
* @param[in] num_indices Number of input indices
* @param[in] filter Filter that can be used to filter out vectors from the merged index
* @param[out] merged_dataset Empty owning dataset handle. Merge first attempts to allocate and
* populate device storage with the same layout as the input indices. For
* an unfiltered merge, if device allocation fails, it falls back to host
* storage and returns a host-backed output index. Keep this dataset alive
* while using \p output_index. A host-backed output index must be updated
* with `cuvsCagraUpdateDataset` before device search.
* @param[in] filter Filter, already applied by the caller while building `merged_dataset`
* @param[in] merged_dataset Caller-owned dataset handle already containing the concatenated (and,
* if `filter` is set, already-filtered) dataset, with the same layout as
* the input indices. Keep this dataset alive while using
* \p output_index. A host-backed dataset must be updated with
* `cuvsCagraUpdateDataset` before device search.
* @param[in] offsets Per-index starting row within `merged_dataset`, as returned by
* `cuvsCagraMergedDatasetOffsets`. Array of `num_indices + 1` int64_t; the last
* entry must equal `merged_dataset`'s row count.
* @param[out] output_index Output handle that will store the merged index.
* Must be initialized using `cuvsCagraIndexCreate` before use.
*/
CUVS_EXPORT cuvsError_t cuvsCagraMerge(cuvsResources_t res,
cuvsCagraIndexParams_t params,
cuvsCagraIndex_t* indices,
size_t num_indices,
cuvsFilter filter,
cuvsDataset_t merged_dataset,
cuvsCagraIndex_t output_index);
CUVS_EXPORT cuvsError_t cuvsCagraMerge_v2(cuvsResources_t res,
cuvsCagraIndexParams_t params,
cuvsCagraIndex_t* indices,
size_t num_indices,
cuvsFilter filter,
cuvsDataset_t merged_dataset,
const int64_t* offsets,
cuvsCagraIndex_t output_index);

/**
* @brief Merge multiple CAGRA indices with explicit merge parameters.
*
* See `cuvsCagraMerge_v2` for the full `merged_dataset`/`offsets` contract.
*
* @param[in] res cuvsResources_t opaque C handle
* @param[in] params cuvsCagraIndexParams_t parameters for the output index
* @param[in] merge_params cuvsCagraMergeParams_t parameters controlling the merge algorithm, or
* NULL to use AUTO defaults
* @param[in] indices Array of input cuvsCagraIndex_t handles to merge
* @param[in] num_indices Number of input indices
* @param[in] filter Filter that can be used to filter out vectors from the merged index
* @param[out] merged_dataset Empty owning dataset handle. Merge first attempts to allocate and
* populate device storage with the same layout as the input indices. For
* an unfiltered merge, AUTO and REBUILD can fall back to host storage if
* device allocation fails; explicit FASTENER reports the allocation
* failure instead. Keep this dataset alive while using `output_index`.
* A host-backed output index must be updated with
* `cuvsCagraUpdateDataset` before device search.
* @param[in] filter Filter, already applied by the caller while building `merged_dataset`
* @param[in] merged_dataset Caller-owned dataset handle already containing the concatenated (and,
* if `filter` is set, already-filtered) dataset. Keep this dataset alive
* while using `output_index`. A host-backed dataset must be updated with
* `cuvsCagraUpdateDataset` before device search.
* @param[in] offsets Per-index starting row within `merged_dataset`, as returned by
* `cuvsCagraMergedDatasetOffsets`. Array of `num_indices + 1` int64_t.
* @param[out] output_index Output handle initialized with `cuvsCagraIndexCreate`
*/
CUVS_EXPORT cuvsError_t cuvsCagraMergeWithParams_v2(cuvsResources_t res,
cuvsCagraIndexParams_t params,
cuvsCagraMergeParams_t merge_params,
cuvsCagraIndex_t* indices,
size_t num_indices,
cuvsFilter filter,
cuvsDataset_t merged_dataset,
const int64_t* offsets,
cuvsCagraIndex_t output_index);

/**
* @brief Merge multiple CAGRA indices into a single CAGRA index (deprecated).
*
* @deprecated Use `cuvsCagraMerge_v2` instead. This function internally allocates and
* populates the merged dataset buffer, hiding memory allocation from the caller.
* `cuvsCagraMerge_v2` requires the caller to pre-allocate and populate
* `merged_dataset`, matching the `cuvsCagraExtend` contract, and gives the caller
* control over peak GPU memory. Will be removed in release 27.02.
*
* @param[in] res cuvsResources_t opaque C handle
* @param[in] params cuvsCagraIndexParams_t parameters for the output index
* @param[in] indices Array of input cuvsCagraIndex_t handles to merge
* @param[in] num_indices Number of input indices
* @param[in] filter Filter applied internally to select which rows to retain
* @param[out] merged_dataset Empty owning dataset handle (created with `cuvsDatasetCreate`).
* Populated internally with the merged dataset; keep it alive while
* using \p output_index and destroy it with `cuvsDatasetDestroy`.
* @param[out] output_index Output handle initialized with `cuvsCagraIndexCreate`
*/
CUVS_EXPORT __attribute__((deprecated(
"Use cuvsCagraMerge_v2 instead; will be removed in 27.02"))) cuvsError_t
cuvsCagraMerge(cuvsResources_t res,
cuvsCagraIndexParams_t params,
cuvsCagraIndex_t* indices,
size_t num_indices,
cuvsFilter filter,
cuvsDataset_t merged_dataset,
cuvsCagraIndex_t output_index);

/**
* @brief Merge multiple CAGRA indices with explicit merge parameters (deprecated).
*
* @deprecated Use `cuvsCagraMergeWithParams_v2` instead. Will be removed in release 27.02.
*
* @param[in] res cuvsResources_t opaque C handle
* @param[in] params cuvsCagraIndexParams_t parameters for the output index
* @param[in] merge_params cuvsCagraMergeParams_t parameters controlling the merge algorithm, or
* NULL to use AUTO defaults
* @param[in] indices Array of input cuvsCagraIndex_t handles to merge
* @param[in] num_indices Number of input indices
* @param[in] filter Filter applied internally to select which rows to retain
* @param[out] merged_dataset Empty owning dataset handle (created with `cuvsDatasetCreate`).
* Populated internally with the merged dataset; keep it alive while
* using `output_index` and destroy it with `cuvsDatasetDestroy`.
* @param[out] output_index Output handle initialized with `cuvsCagraIndexCreate`
*/
CUVS_EXPORT cuvsError_t cuvsCagraMergeWithParams(cuvsResources_t res,
cuvsCagraIndexParams_t params,
cuvsCagraMergeParams_t merge_params,
cuvsCagraIndex_t* indices,
size_t num_indices,
cuvsFilter filter,
cuvsDataset_t merged_dataset,
cuvsCagraIndex_t output_index);
CUVS_EXPORT __attribute__((deprecated(
"Use cuvsCagraMergeWithParams_v2 instead; will be removed in 27.02"))) cuvsError_t
cuvsCagraMergeWithParams(cuvsResources_t res,
cuvsCagraIndexParams_t params,
cuvsCagraMergeParams_t merge_params,
cuvsCagraIndex_t* indices,
size_t num_indices,
cuvsFilter filter,
cuvsDataset_t merged_dataset,
cuvsCagraIndex_t output_index);

/**
* @}
Expand Down
Loading
Loading