Repeating unlearning on one trained model measures variation conditional on that model. It does not reveal how the result changes when the original model is trained again. When training contributes to variability, extra unlearning runs cannot generally replace independent training runs. Lanyon et al. explain this through a training/unlearning variance decomposition and give guidance on allocating compute between the two.
SUPREME makes this experimental design practical: train I original models, run J unlearning repetitions per model, and optionally K evaluation repetitions per unlearned model. Separate stage seeds support investigating where variation arises. An extensible Python API supports comparisons against retraining, with execution from one GPU to a cluster.
Seed design and variance analysis · Research on training seeds
Existing Table 1 results on Pins Face Recognition, random-sample unlearning (forget 0.1%). Accuracy differences are unlearned minus retrained, in percentage points. Bars show one standard deviation across ten matched-seed pipelines (J = K = 1), combining variation across stages. These experiments used one GPU.
Interactive viewer · Download the tables · Paper · Presentation
Browse the published tables with Python 3.9 or later and its standard library:
git clone https://github.com/pedroandreou/supreme-unlearning.git
cd supreme-unlearning
python3 examples/paper_results.pyThe results guide covers downloads, measurement definitions and exporting an offline viewer.
SUPREME combines multi-seed image-unlearning evaluation with multi-GPU execution and configurable numerical precision.
| Framework | Domain in the comparison | Multi-seed | Multi-GPU | Multi-precision |
|---|---|---|---|---|
| OpenUnlearning | LLMs | Not shown | Yes | Yes |
| MUBox | Image classification | Not shown | Not shown | Not shown |
| ERASURE | Image classification | Yes | Not shown | Not shown |
| Deep Unlearn | Image classification | Yes | Not shown | Not shown |
| SUPREME | Image classification | Yes | Yes | Yes |
Feature definitions and sources.
| Component | Included |
|---|---|
| Datasets | CIFAR-10, CIFAR-20, CIFAR-100, Pins Face Recognition, Caltech-101 |
| Models | ResNet18, Vision Transformer |
| Methods | FT, Bad Teacher, Random Labels, UNSIR, SSD, LFSSD, SSD-Det, LFSSD-Det, ASSD, SCRUB, JIT |
| Reference models | Retrain and Original |
| Scenarios | Full-class, subclass and random-sample unlearning |
| Evaluation | Accuracy, membership inference, model distances and resource measurements |
Full component and hardware reference.
Install the Python library:
pip install supreme-unlearningFor a complete train → unlearn → evaluate run, follow the experiment quickstart. It covers environment setup, credentials and a small example.
The framework uses the paper's pinned dependency stack. Check the platform requirements and security guidance before loading external models or checkpoints.
Register components from your own Python package:
import supreme
# Replace the module path with your implementation.
supreme.register_unlearning_method("mymethod", "your_package.your_method")The library guide describes the public API and pipeline calls; the extension guide covers component interfaces.
| I want to… | Start here |
|---|---|
| Run local or SLURM experiments | Experiment guide |
| Reproduce the paper | Reproduction guide |
| Study variation across stages | Seed protocols |
| Add a method, metric, model or dataset | Custom-component notebook |
All documentation, including notation, implementation details, logging, tooling and maintainer workflows.
Bug reports, new components and documentation contributions are welcome. Open an issue, read the contributing guide, or share a method and its results.
If you use SUPREME in your research, please cite our paper and the original papers for the methods you use. See method credits and citations.
@misc{supreme2026,
title = {SUPREME: A Multi-GPU Framework for Reproducible Image Unlearning Method Evaluation},
author = {Petros Andreou, Jamie Lanyon, Axel Finke, Georgina Cosma},
year = {2026},
eprint = {2606.00380},
archivePrefix = {arXiv},
primaryClass = {cs.LG},
url = {https://arxiv.org/abs/2606.00380}
}This work was conducted at Loughborough University.
SUPREME builds on SSD, Bad Teacher and other open-source unlearning research. We thank their authors; full credits and citation guidance are available for each method.
This project is licensed under the MIT License.
If SUPREME is useful for your research, star the repository to keep it handy.
