Skip to content

examples: train in place on a deep-chunked cloud archive (dynamical.org NOAA GFS) - #92

Open
emfdavid wants to merge 1 commit into
mainfrom
examples/dynamical-downscaling
Open

emfdavid wants to merge 1 commit into
mainfrom
examples/dynamical-downscaling

Conversation

@emfdavid

Copy link
Copy Markdown
Owner

Summary

Adds examples/dynamical/ — spatial downscaling of surface solar irradiance on
dynamical.org's public NOAA GFS analysis, an Icechunk repository whose
stored chunk holds 1440 consecutive hourly global fields. It is the deep-chunk end of the
example set: one fetch and one decode serve 1440 samples, where a per-sample __getitem__ would
re-read the same object every time.

It exists because pointing the README quickstart at one of these archives fails three ways
before a byte moves, and none of them raise:

  1. the stores are Icechunk, so obstore_store() is the wrong constructor
  2. variables= stops being optional at 25–146 arrays per store (and sample_axis= then raises
    on the 1-D coordinate arrays)
  3. the quickstart's block_chunks=16 is sized to a 63 MiB chunk; NOAA GFS analysis has a
    6.87 GiB one, and three variables at those defaults asks print_summary() for 860 GiB
    of estimated peak — which it truthfully reports

The example reads one variable on purpose. A stored chunk is 878.9 MiB and assembles to
6.87 GiB resident; the floor is two blocks of them, so each variable costs ~13.7 GiB whatever the
batch size, and a shrinking chunk_transform cannot buy it back (slot_charge_bytes takes
max(source tiles, assembled output), because the pool holds the source tiles until the fill
completes — so a Coarsen(4) leaves residency unchanged and raises peak). Picking a smaller
store would have hidden the one thing this geometry is here to teach.

Task is downscaling over a European window; baseline is bilinear upsampling — what the coarse
field gives you with no model, and a data fingerprint besides, since its RMSE depends only on the
bytes the loader handed over. Follows the microscopy/ shape: data.py framework-neutral,
train_torch.py the only framework file, offline synthetic --source so it runs with no network
or credentials.

Docs, examples/README.md and a CHANGELOG bullet included. No engine code is touched — this is
examples, tests and docs only.

For reviewers

Most valuable second look:

  • The residency claim in the docs and CHANGELOG. I read it out of slot_charge_bytes rather
    than measuring a shrinking transform's peak RSS directly; the describe() numbers agree with
    the reading, but the arithmetic is worth a check.
  • region_index() — the European window wraps the prime meridian, so it returns wrapping
    column indices rather than a slice. test_region_crop_wraps_the_prime_meridian asserts exactly
    one discontinuity; a plain slice would silently take the long way round the globe.
  • Whether the synthetic store earns its construction. Its sub-grid structure is a
    deterministic function of the smooth field, so the CNN beats bilinear by construction, which is
    what makes the test fast and deterministic. That is stated in the docstring rather than implied.

Confident in: the geometry assertions, and that the real-archive path runs (log below).

Verification

All green locally on this branch:

  • uv run ruff check src tests bench examples ✅
  • uv run ruff format --check src tests bench examples ✅
  • uv run mypy src bench examples ✅
  • uv run pytest -q → 673 passed, 9 skipped ✅
  • uv run --extra docs mkdocs build --strict ✅

Example runs (these are demonstration runs, not a performance claim — no baseline engine is
being compared against):

run result
--source synthetic (offline, 8 epochs) model 28.6 vs bilinear 35.4 W/m²
--source gfs (2 epochs, default 4-chunk range) model 22.65 vs bilinear 25.10 W/m²

The GFS run, on an 8 vCPU / 31 GiB n2-standard-8 in GCP reading s3://dynamical-noaa-gfs in
us-west-2 (cross-cloud, so pessimistic): 90 batches in 102.7 s cold at 0% chunk hit, then
93.0 s at 100% chunk hit with zero fetches and zero decodes. print_summary() estimated
19.78 GiB of peak on the defaults, which is what it took.

Author attestation

  • I have reviewed every change in this PR, I can explain why each one is correct, and I
    have verified the claims made in this description.

Left unchecked deliberately — this was AI-assisted and the attestation is the human author's to
tick.

Checklist

  • Tests added or updated — tests/test_dynamical.py, 5 tests, synthetic only (11 s)
  • ruff check, ruff format --check, mypy, pytest -q and mkdocs build --strict green
    locally
  • Docstrings for the new public surface (all of examples/dynamical/data.py)
  • User-facing behavior documented — docs/examples.md (table row + section) and
    examples/README.md
  • A bullet added under ## Unreleased in CHANGELOG.md
  • No load-bearing invariant broken — examples/tests/docs only, no engine change; Batch
    stays numpy, no dask, no reshard, no xr.DataArray
  • Does not touch ChunkPool, the scheduler, or cross-thread readiness
  • No performance claim made — the numbers above are demonstration runs with their hardware
    and store stated, not a comparison against another engine

🤖 Generated with Claude Code

https://claude.ai/code/session_016eLANSQjFGbLQprE7LbyTJ

dynamical.org publishes ML-shaped weather Zarr with no published path into a
training loop, and pointing the README quickstart at one fails three ways
before a byte moves: the stores are Icechunk rather than plain Zarr,
`variables=` stops being optional at 25 to 146 arrays per store, and the
quickstart's `block_chunks=16` is sized to a 63 MiB chunk where NOAA GFS
analysis has a 6.87 GiB one.

The task is spatial downscaling of surface solar irradiance over a European
window, with bilinear upsampling as the model-free baseline. It reads one
variable because the geometry prices the second one at another 13.7 GiB of
residency, and the example says so rather than picking a small store that
would have hidden it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016eLANSQjFGbLQprE7LbyTJ

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant