Skip to content

Add Float32VectorSparse with fromMap constructor - #187

Closed
gyrdym wants to merge 16 commits into
masterfrom
cursor/float32-vector-sparse-f5ac
Closed

gyrdym wants to merge 16 commits into
masterfrom
cursor/float32-vector-sparse-f5ac

Conversation

@gyrdym

@gyrdym gyrdym commented Jul 26, 2026 •

Copy link
Copy Markdown
Owner

Summary

Adds Float32VectorSparse, a float32 sparse vector implementation that satisfies the full Vector interface.

Details

  • Constructor: Float32VectorSparse.fromMap(Map<int, num> source, {required int length})
  • Stores sorted non-zero index/value pairs; missing positions are 0.0
  • Algebraic == / hashCode by vector values (works across Vector implementations from the sparse side)
  • Sparse-friendly paths for [], sum, dot, norm, min/max, abs, scalar *//, sparse *, set, subvector
  • Remaining operations currently fall back to a dense Vector for correctness
  • Basic unit tests for construction and core behavior

Notes

Branch rebased onto rewritten master.

Open in Web Open in Cursor 

Introduce a float32 sparse vector that stores only non-zero entries,
fully satisfies the Vector interface, and includes basic unit tests.

Co-authored-by: Ilia Gyrdymov <gyrdym@users.noreply.github.com>
@cursor
cursor Bot force-pushed the cursor/float32-vector-sparse-f5ac branch from a839996 to d439048 Compare July 27, 2026 21:22
cursoragent and others added 15 commits July 28, 2026 22:38
Co-authored-by: Ilia Gyrdymov <gyrdym@users.noreply.github.com>
Compare against any Vector by length and elements, and hash the full
logical value sequence so equal sparse vectors share hashCode.

Co-authored-by: Ilia Gyrdymov <gyrdym@users.noreply.github.com>
Hash length and stored index/value pairs with a stronger mixing
finalizer instead of scanning the full dense logical vector.

Co-authored-by: Ilia Gyrdymov <gyrdym@users.noreply.github.com>
Move sparse vector hashing helpers to lib/src/common/hash with dedicated
unit tests for each function.

Co-authored-by: Ilia Gyrdymov <gyrdym@users.noreply.github.com>
Keep Float32VectorSparse.sqrt on the sparse path when nnz is below
half the length; otherwise densify for SIMD-friendly evaluation.

Co-authored-by: Ilia Gyrdymov <gyrdym@users.noreply.github.com>
Mirror sqrt: keep the sparse path below half density, otherwise densify
for SIMD-friendly abs.

Co-authored-by: Ilia Gyrdymov <gyrdym@users.noreply.github.com>
Co-authored-by: Ilia Gyrdymov <gyrdym@users.noreply.github.com>
Keep the sparse path for positive exponents when clearly sparse;
otherwise densify. Zero/negative exponents keep their existing behavior.

Co-authored-by: Ilia Gyrdymov <gyrdym@users.noreply.github.com>
Use the sparse path for scalar *, scalar /, and vector * when clearly
sparse; otherwise densify for SIMD-friendly evaluation.

Co-authored-by: Ilia Gyrdymov <gyrdym@users.noreply.github.com>
Introduce a fill value for missing indices so exp can map fill 0 -> 1
and stay sparse. Aggregations and element access honor fill; structural
multiply densifies when fill is non-zero.

Co-authored-by: Ilia Gyrdymov <gyrdym@users.noreply.github.com>
Return Float32VectorSparse to classic zero-fill semantics. exp goes back
to a dense path, which better matches on-device text workloads.

Co-authored-by: Ilia Gyrdymov <gyrdym@users.noreply.github.com>
Remove the nnz heuristic from * and scalar / so text-style sparse
vectors stay compressed for the common on-device workloads.

Co-authored-by: Ilia Gyrdymov <gyrdym@users.noreply.github.com>
Exercise sparse vectors end-to-end without mocks: large-vocab BoW
storage, cosine ranking against a dense baseline, feature intersection,
and live set updates.

Co-authored-by: Ilia Gyrdymov <gyrdym@users.noreply.github.com>
Co-authored-by: Ilia Gyrdymov <gyrdym@users.noreply.github.com>
Co-authored-by: Ilia Gyrdymov <gyrdym@users.noreply.github.com>
@gyrdym gyrdym closed this Sep 27, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants