This repository contains the checkpoint-generation and graph-analysis code for the zebrafish proofreading study, published in Connectome quality converges predictably to reveal optimal stopping points during proofreading. The active workflow now follows a science-repo layout:
data/raw/: external inputs such aschanges.tsv, connectome CSVs, and contactome CSVsdata/provided/: stable project-supplied inputsdata/generated/: derived intermediates such as component-graph snapshots and invariant JSONsanalysis/: workflow entrypointsnotebooks/: legacy notebooks that are useful to keep aroundsrc/zfish/: shared Python packagesrc/hemibrain/: standalone slurm scripts for Hemibrain analysisresults/data/: final summary tablesresults/figures/: final figuresSnakefile: reproducible workflow wiring
data/ remains ignored in git. Code and results/ are the tracked outputs.
uv sync --group devThis workflow converts an Ariadne proofreading changelog into a series of segmentation checkpoints that can be visualized or analyzed. The process consists of four stages:
- Download the proofreading changelog
- Reconstruct intermediate agglomeration graphs
- Generate supervoxel remapping tables
- Apply each remapping to the original supervoxel segmentation to create checkpoint volumes
The default directory structure is
data/
├── raw/
│ └── changes.tsv
└── generated/
├── component-graphs/
└── remaps/
Each stage uses the output of the previous stage.
Download the proofreading edit history from your Ariadne deployment.
wget https://<ariadne-domain>/api/changes.tsv \
-O data/raw/changes.tsvThis file records every merge and split operation performed during proofreading.
Reconstruct the agglomeration state after every N proofreading edits.
uv run python manage.py changelog subsample| Argument | Default | Description |
|---|---|---|
--changes-file |
data/raw/changes.tsv |
Input Ariadne changelog |
--output-dir |
data/generated/component-graphs |
Directory for graph snapshots |
--step-size |
5000 |
Number of proofreading edits between checkpoints |
--max-changes |
Process entire changelog | Stop after N edits |
--shuffle-count |
0 |
Generate randomized proofreading orders |
Create lookup tables mapping every supervoxel ID to its agglomerated object ID for each checkpoint.
uv run python manage.py remap generate| Argument | Default |
|---|---|
--component-graph-dir |
data/generated/component-graphs |
--output-dir |
data/generated/remaps |
Apply each remapping dictionary to the original supervoxel segmentation.
uv run python manage.py checkpoint upload \
--supervoxels-path precomputed://<path-to-supervoxels> \
--checkpoint-prefix precomputed://<output-prefix>The result is a sequence of segmentation checkpoints representing the dataset at different stages of proofreading.
| Argument | Default |
|---|---|
--supervoxels-path |
(required) CloudVolume containing the original supervoxel segmentation |
--checkpoint-prefix |
(required) Prefix used when writing checkpoint segmentations |
--component-graph-dir |
data/generated/component-graphs |
--remap-dir |
data/generated/remaps |
--mip |
0 |
--jobs |
1 |
--block-size |
(512, 512, 64) |
--chunk-size |
(512, 512, 64) |
--skip |
0 |
--shuffle-id |
None |
--randomize |
Enabled (disable with --no-randomize) |
--progress |
Enabled (disable with --no-progress) |
Put one ordered connectome series under data/raw/connectomes/original_order/ and one ordered contactome series under data/raw/contactomes/original_order/. The loaders also accept the legacy data/connectomes and data/contactomes paths for one migration cycle.
Run the full analysis with Snakemake:
uv run snakemake --cores 1This produces:
data/generated/connectome/invariants.jsondata/generated/contactome/invariants.jsonresults/data/*-convergence-summary.mdresults/data/*-convergence-summary.texresults/data/*-convergence-stats.jsonresults/figures/*-convergence/results/data/*-projection-summary.mdresults/data/*-projection-summary.texresults/data/*-projection-summary.csvresults/data/*-projection-stats.jsonresults/figures/*-projection-milestones/
The default metric profile is fast. It prefers scalable or approximate graph metrics and avoids exhaustive computations such as exact diameter in routine runs.
You can run workflow steps directly while developing:
uv run python analysis/invariants/run.py \
--dataset connectome \
--input-dir data/raw/connectomes/original_order \
--output-json data/generated/connectome/invariants.json \
--metric-profile fast
uv run python analysis/convergence/run.py \
--dataset connectome \
--input-json data/generated/connectome/invariants.json \
--summary-md results/data/connectome-convergence-summary.md \
--summary-tex results/data/connectome-convergence-summary.tex \
--stats-json results/data/connectome-convergence-stats.json \
--figure-dir results/figures/connectome-convergence
uv run python analysis/projections/run.py \
--dataset connectome \
--input-json data/generated/connectome/invariants.json \
--convergence-stats results/data/connectome-convergence-stats.json \
--summary-md results/data/connectome-projection-summary.md \
--summary-tex results/data/connectome-projection-summary.tex \
--summary-csv results/data/connectome-projection-summary.csv \
--stats-json results/data/connectome-projection-stats.json \
--figure-dir results/figures/connectome-projection-milestonesFor questions or collaboration inquiries, please email hannah.martinez@jhuapl.edu or jordan.matelsky@jhuapl.edu.
This software was created by the Johns Hopkins University Applied Physics Laboratory, with funding supported by the NIH BRAIN Initiative under grant no. R24MH114785.
The views, opinions, and/or findings expressed are those of the author(s) and should not be interpreted as representing the official views or policies of the NIH.
© 2026 The Johns Hopkins University Applied Physics Laboratory LLC