Skip to content

Read and write icechunk sessions through the zarr-python Store instead of linking the icechunk crate #196

Description

@kylebarron

Note

This issue was written by Claude (Claude Code, Opus 5), not by @kylebarron.

Problem

Writes through an icechunk Session are lost (#187). This is the concern raised in #107.

zarrista cannot share a Rust object with the separately compiled icechunk Python package. So PyAsyncIcechunkStore serializes the Python session (the private _session.as_bytes()) and rebuilds a second icechunk::session::Session in Rust (src/storage/async.rs). All writes go into the change set of that copy. Nothing copies the change set back, so:

  • session.has_uncommitted_changes stays False, and session.commit() raises SessionStateError: no changes made to the session.
  • Each extraction builds a new copy. AsyncArray.open(session, "/arr") after builder.create_async(session, "/arr") fails with array metadata is missing.
  • A ForkSession is rejected with TypeError, because is_icechunk_session checks for the exact class name Session.

Merging the Rust copy back (serialize it, wrap it in a ForkSession, call session.merge) works in a quick test, but it adds more to the private-API, two-instance design instead of removing it.

Proposal

Read and write through the public zarr-python Store (session.store) instead of linking the icechunk crate. Implement this as a generic Rust adapter over any zarr-python async Store, not an icechunk-specific one.

Benefits:

  • One session and one change set. Writes appear in session, and commit works.
  • No private API. No reconstructed second icechunk instance.
  • No version coupling. The icechunk major-version check and the icechunk>=2 / Python 3.12 split in pyproject.toml are no longer needed.
  • icechunk.in_memory_storage() sessions work.
  • The icechunk and zarrs_icechunk crates leave the async feature. This reduces compile time and wheel size.
  • Other zarr-python stores also work.

Costs and risks:

  • Each store operation goes through Python: take the GIL, schedule a coroutine on the asyncio event loop, and wait. Decoding stays parallel in Rust, but all I/O dispatch goes through one event loop. This can matter for many small keys, such as sharded partial reads.
  • The public icechunk store copies data on get (prototype.buffer.from_bytes) and on set (value.to_bytes()).
  • zarrs calls the store from tokio threads. The adapter must capture TaskLocals when a call starts and convert Python coroutines with pyo3_async_runtimes::into_future_with_locals.

Plan

  1. Prototype the adapter with get, get_partial_values, set, delete, list, and list_dir.
  2. Benchmark reads against the current linked-crate path. Use local storage, a full-array read, and a sharded partial read.
  3. If the overhead is acceptable, replace PyAsyncIcechunkStore with the adapter for reads and writes, and remove the icechunk and zarrs_icechunk dependencies.
  4. If read overhead is too large, consider the linked crate for read-only sessions only (a copy of a read-only session is safe) and the adapter for writable sessions. Avoid this unless the numbers require it, because it keeps two code paths.
  5. Add write tests to tests/test_icechunk.py: create and write with zarrista, commit, and read back. Today the tests write with zarr-python and only read with zarrista, so they did not catch Can't write to icechunk store #187.

Closes #187. Resolves #107.


🤖 Written by Claude Code

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions