This repository is the hub that ties together the work I carried out during my end-of-studies internship at the CosmoStat laboratory (CEA Paris-Saclay), from April 2026 to September 2026, supervised by Samuel Farrens and Emma Ayçoberry.
The goal was to build an end-to-end pipeline for generative modelling of galaxy images: assemble a training dataset from real Euclid observations, train a deep generative model on an HPC cluster, and check how faithfully the generated galaxies reproduce the morphology of real ones.
The full write-up is available here:
Internship_report_final.pdf.
The work is split across three standalone repositories, one per stage of the pipeline.
1. Dataset construction — Euclid-Q1-postage-stamps
Turns Euclid Q1 VIS calibrated imaging into a machine-learning-ready dataset of 64×64 postage stamps of isolated sources, each paired with its noise map, a bad-pixel mask, and a locally interpolated PSF (optionally a residual PSF kernel). The pipeline has three stages:
- Acquisition — query the Euclid Science Archive; download the calibrated science frames, their background frames, the global VIS PSF model, and the cross-matched source catalogues; slice everything into per-quadrant FITS files.
- Build — for every selected, isolated catalogue source (cuts on flux,
point_like_prob,spurious_prob, and bad-pixel fraction), cut a background-subtracted stamp with its noise and mask, interpolate the PSF at the source position, and collect the records into a Hugging Facedatasets.Dataset. - Output — save the dataset locally and/or push it to the Hugging Face Hub.
2. Model training — Train-AE
Trains the two-stage generative model on the 64×64 galaxy postage stamps, run on the Jean Zay supercomputer (IDRIS/CNRS):
- A convolutional autoencoder that reconstructs galaxy images, convolving
its output with the instrument PSF (via
jax-galsim) before comparing it to the observed, PSF-convolved image. Trained in a full-PSF variant (with a total-variation term to tame the ill-posed deconvolution) and a partial-PSF variant that avoids the non-physical regularisation. - A normalising flow fitted on the frozen autoencoder's latent space, so new galaxies can be sampled by drawing latent codes from the flow and decoding them.
Written in JAX/Equinox, training logged to Weights & Biases. The modelling code builds on prior work by Benjamin Rémy (former CosmoStat PhD student).
3. Result verification — galaxy-morphometrics
A standalone toolkit that checks the generated images by comparing the
morphology of real, reconstructed, and flow-sampled galaxies. It reproduces
the validation plots of Lanusse et al. (2020) /
deep_galaxy_models,
computing per-object statistics — HSM adaptive moments (size, ellipticity,
rho4), CAS (concentration, asymmetry), Gini/M20, and MID — and rendering their
distributions with one curve per dataset (real, reconstruction, flow prior).
Train-AE also carries a complementary PQMass check that generated samples match
the held-out data distribution in both latent and image space.
Euclid Q1 VIS calibrated imaging
│
▼
[Euclid-Q1-postage-stamps] → 64×64 stamps (+ noise, mask, PSF) on the HF Hub
│
▼
[Train-AE] → autoencoder + latent normalising flow, trained on Jean Zay
│
▼
[galaxy-morphometrics] → morphology of real vs. reconstructed vs. generated galaxies
| Laboratory | CosmoStat, CEA Paris-Saclay |
| Period | April 2026 – September 2026 |
| Supervisors | Samuel Farrens, Emma Ayçoberry |
| Compute | Jean Zay supercomputer (IDRIS/CNRS) |
| Data | Euclid Q1 release, VIS instrument |