A small PyTorch project that implements a fully connected Autoencoder from scratch and trains it on the MNIST handwritten-digit dataset.
The project focuses on understanding how Autoencoders learn compressed representations of images through an encoder–decoder architecture. The notebook covers the complete workflow, from loading MNIST and training the model to reconstructing images and visualizing the learned latent space.
An Autoencoder learns to reconstruct its input:
Input Image
│
▼
Encoder
│
▼
Latent Representation
│
▼
Decoder
│
▼
Reconstructed Image
Instead of learning to classify digits, the model learns a compressed representation that preserves the information necessary to reconstruct the original image.
For this project:
784 → 256 → 128 → 256 → 784
The 128-dimensional latent representation acts as the bottleneck of the network.
.
├── pyproject.toml
├── requirements.txt
├── src/
│ └── autoencoder/
│ ├── __init__.py
│ ├── Autoencoder_implimentation.ipynb
│ └── data/
│ └── MNIST/
└── README.md
The main implementation is contained in:
src/autoencoder/Autoencoder_implimentation.ipynb
The notebook is the executable entry point for the project. The Python package module is currently a minimal placeholder.
- Python 3.14 or newer
- PyTorch
- torchvision
- NumPy
- Matplotlib
- scikit-learn
- Jupyter/IPython kernel
- CUDA-capable GPU (optional)
The notebook automatically selects CUDA when a compatible GPU is available and otherwise falls back to the CPU.
git clone <repository-url>
cd AutoencoderUsing Python's built-in venv:
py -3.14 -m venv .venvActivate it:
.\.venv\Scripts\Activate.ps1python -m pip install --upgrade pip
python -m pip install -r requirements.txtTo install the package in editable mode:
python -m pip install -e .The project also contains a pyproject.toml configured for uv. Install uv first if it is not already available.
If uv is installed:
uv venv
.\.venv\Scripts\Activate.ps1
uv pip install -r requirements.txt
uv pip install -e .Open the notebook at:
src/autoencoder/Autoencoder_implimentation.ipynb
in VS Code or Jupyter Notebook.
Then:
- Select the project's
.venvas the notebook kernel. - Run the cells from top to bottom.
- Allow
torchvisionto download MNIST automatically. - Inspect the reconstruction results.
- Explore the latent-space visualizations.
The dataset does not need to be downloaded manually.
The dataset cell uses:
root="./data"for the MNIST dataset. This path is relative to the notebook kernel's current working directory, not to the notebook file itself. When the notebook is run from the repository root, the dataset will be stored under:
data/MNIST/
If the notebook is launched with a different working directory, MNIST will be downloaded to that directory's data/MNIST/ folder.
The autoencoder console command defined in pyproject.toml currently calls a placeholder main() function and is not the training entry point. Run the notebook for the actual implementation.
The notebook follows a progressive learning workflow:
Import the required libraries and automatically select CPU or CUDA.
Load MNIST, convert images to tensors, and create batches using a PyTorch DataLoader.
MNIST images have the shape:
1 × 28 × 28
Because the model uses fully connected layers, each image is flattened into:
784 values
Define separate encoder and decoder networks using PyTorch's nn.Sequential and nn.Linear layers.
Train the network using:
- MSE reconstruction loss
- Adam optimizer
- Learning rate:
0.001 - Batch size:
128 - Epochs:
20
Compare original MNIST images with the images reconstructed by the decoder.
Pass the complete training dataset through the encoder and collect the resulting latent vectors.
Analyze the learned representation directly when possible or use PCA when the latent space has more than three dimensions.
Each MNIST image is a grayscale image of size:
28 × 28
The image is flattened before being passed to the fully connected network.
INPUT
1 × 28 × 28
│
▼
Flatten
│
▼
784 values
│
▼
Linear + ReLU
│
▼
256 values
│
▼
Linear + ReLU
│
▼
Latent Space
128 values
│
▼
Linear + ReLU
│
▼
256 values
│
▼
Linear + Sigmoid
│
▼
784 values
│
▼
Reshape
│
▼
28 × 28 IMAGE
The encoder performs the compression:
784 → 256 → 128
The bottleneck contains:
128 values
This representation is a compressed version of the original 784-dimensional image.
The decoder reconstructs the image:
128 → 256 → 784
The final Sigmoid activation produces values in the [0, 1] range, matching the normalized pixel values produced by transforms.ToTensor().
The Autoencoder does not require the MNIST digit labels during training.
Instead, the input itself becomes the target:
┌───────────────┐
│ │
▼ │
x → Encoder → z → Decoder → x̂
│
▼
Reconstruction
Loss
Where:
x= original imagez= latent representationx̂= reconstructed image
The training objective is to make:
x̂ ≈ x
The project uses Mean Squared Error (MSE):
MSE(x, x̂) = mean((x - x̂)²)
The loss compares every pixel of the reconstructed image with the corresponding pixel in the original image.
During training, the optimizer updates the network parameters to minimize this reconstruction error.
The model uses the Adam optimizer:
optimizer = optim.Adam(
model.parameters(),
lr=0.001
)Training configuration:
| Parameter | Value |
|---|---|
| Batch Size | 128 |
| Epochs | 20 |
| Learning Rate | 0.001 |
| Optimizer | Adam |
| Loss | MSE |
| Latent Dimension | 128 |
| Hidden Dimension | 256 |
After training, the encoder is used to generate a latent representation for every MNIST image.
With the default architecture:
60,000 images × 128 latent dimensions
The MNIST labels are collected only for visualization.
They are not used to train the Autoencoder.
This allows us to investigate whether visually similar digits naturally occupy nearby regions of the learned latent space.
The notebook automatically chooses a visualization based on the latent dimensionality.
For one-dimensional representations:
- Class-specific histograms
- Latent-value strip plot
For two-dimensional representations:
z₁ vs z₂
is plotted directly.
For three-dimensional representations, a 3D scatter plot is generated.
For latent spaces with more than three dimensions, Principal Component Analysis (PCA) is used.
For the default 128-dimensional representation:
128D latent space
│
▼
PCA
│
▼
2D space
The notebook also reports the explained variance ratio of the selected principal components.
The notebook produces two primary types of results.
Original images are compared against their reconstructed versions:
Original Reconstruction
───────── ───────────────
digit → generated
image reconstruction
This provides a visual indication of how well the Autoencoder has learned to preserve the important information in the input images.
The learned latent representations are projected into a space that can be visualized.
By coloring the points according to their MNIST labels, we can inspect whether different digits form distinguishable regions.
Note: Because the labels are not used during training, any visible grouping is a property of the representation learned through reconstruction rather than supervised classification.
This project provides a practical implementation of several important deep-learning concepts:
- Neural network encoders and decoders
- Representation learning
- Dimensionality reduction
- Bottleneck architectures
- Self-supervised learning
- Reconstruction loss
- Backpropagation
- Adam optimization
- PyTorch
nn.Module - PyTorch
DataLoader - GPU/CPU device management
- Latent-space analysis
- PCA visualization
The primary goal of this project is not simply to train an Autoencoder, but to understand how neural networks can learn useful representations without explicit class labels.
The key idea is:
High-dimensional input
│
▼
Encoder
│
▼
Compact representation
│
▼
Decoder
│
▼
Reconstructed input
The bottleneck forces the network to learn which information is important enough to preserve.
Python
PyTorch
Torchvision
NumPy
Matplotlib
Scikit-learn
Jupyter Notebook