Skip to content

Repository files navigation

Evaluate unlearning beyond a single trained model

Repeating unlearning on one trained model measures variation conditional on that model. It does not reveal how the result changes when the original model is trained again. When training contributes to variability, extra unlearning runs cannot generally replace independent training runs. Lanyon et al. explain this through a training/unlearning variance decomposition and give guidance on allocating compute between the two.

SUPREME makes this experimental design practical: train I original models, run J unlearning repetitions per model, and optionally K evaluation repetitions per unlearned model. Separate stage seeds support investigating where variation arises. An extensible Python API supports comparisons against retraining, with execution from one GPU to a cluster.

Seed design and variance analysis · Research on training seeds

See what the paper reports

Published forget-accuracy differences, showing means and standard deviations across ten seeds

Existing Table 1 results on Pins Face Recognition, random-sample unlearning (forget 0.1%). Accuracy differences are unlearned minus retrained, in percentage points. Bars show one standard deviation across ten matched-seed pipelines (J = K = 1), combining variation across stages. These experiments used one GPU.

Interactive viewer · Download the tables · Paper · Presentation

Try the results example

Browse the published tables with Python 3.9 or later and its standard library:

git clone https://github.com/pedroandreou/supreme-unlearning.git
cd supreme-unlearning
python3 examples/paper_results.py

The results guide covers downloads, measurement definitions and exporting an offline viewer.

Framework comparison

SUPREME combines multi-seed image-unlearning evaluation with multi-GPU execution and configurable numerical precision.

Framework Domain in the comparison Multi-seed Multi-GPU Multi-precision
OpenUnlearning LLMs Not shown Yes Yes
MUBox Image classification Not shown Not shown Not shown
ERASURE Image classification Yes Not shown Not shown
Deep Unlearn Image classification Yes Not shown Not shown
SUPREME Image classification Yes Yes Yes

Feature definitions and sources.

🗃️ Available Components

Component Included
Datasets CIFAR-10, CIFAR-20, CIFAR-100, Pins Face Recognition, Caltech-101
Models ResNet18, Vision Transformer
Methods FT, Bad Teacher, Random Labels, UNSIR, SSD, LFSSD, SSD-Det, LFSSD-Det, ASSD, SCRUB, JIT
Reference models Retrain and Original
Scenarios Full-class, subclass and random-sample unlearning
Evaluation Accuracy, membership inference, model distances and resource measurements

Full component and hardware reference.

⚡ Quickstart

Install the Python library:

pip install supreme-unlearning

For a complete train → unlearn → evaluate run, follow the experiment quickstart. It covers environment setup, credentials and a small example.

The framework uses the paper's pinned dependency stack. Check the platform requirements and security guidance before loading external models or checkpoints.

📦 SUPREME as a Library

Register components from your own Python package:

import supreme

# Replace the module path with your implementation.
supreme.register_unlearning_method("mymethod", "your_package.your_method")

The library guide describes the public API and pipeline calls; the extension guide covers component interfaces.

📚 Documentation

I want to… Start here
Run local or SLURM experiments Experiment guide
Reproduce the paper Reproduction guide
Study variation across stages Seed protocols
Add a method, metric, model or dataset Custom-component notebook

All documentation, including notation, implementation details, logging, tooling and maintainer workflows.

🤝 Contributing

Bug reports, new components and documentation contributions are welcome. Open an issue, read the contributing guide, or share a method and its results.

📝 Citing this work

If you use SUPREME in your research, please cite our paper and the original papers for the methods you use. See method credits and citations.

@misc{supreme2026,
  title  = {SUPREME: A Multi-GPU Framework for Reproducible Image Unlearning Method Evaluation},
  author = {Petros Andreou, Jamie Lanyon, Axel Finke, Georgina Cosma},
  year   = {2026},
  eprint = {2606.00380},
  archivePrefix = {arXiv},
  primaryClass = {cs.LG},
  url    = {https://arxiv.org/abs/2606.00380}
}

This work was conducted at Loughborough University.

🙏 Acknowledgements

SUPREME builds on SSD, Bad Teacher and other open-source unlearning research. We thank their authors; full credits and citation guidance are available for each method.

📄 License

This project is licensed under the MIT License.

If SUPREME is useful for your research, star the repository to keep it handy.

About

Multi-GPU framework for reproducible image-unlearning evaluation, with independent training, unlearning and evaluation seeds and extensible datasets, models, methods and metrics. Published at WIPE-OUT 2 (ECML-PKDD 2026).

Topics

Resources

Contributing

Security policy

Stars

6 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages