The goal of this project is to reconstruct missing regions in images using machine learning techniques. The reconstructed images are compared with the original ground-truth images using quantitative metrics such as Mean Squared Error (MSE) and Mean Absolute Error (MAE), as well as visual examples.
- Dataset: Curated using 10 selected classes from ImageNet: butterfly, panda, parrot, pomeranian, goldfish, elephant, monkey, Persian cat, penguin, and red panda.
- Image Dimensions: Each image is resized to (224 imes 224) pixels.
- Masked Regions: Blacked-out regions of random sizes ranging from (30 imes 30) to (60 imes 60) pixels simulate missing data. The positions of the masks are randomized.
- Dataset Splits:
- Training: 70%
- Validation: 10%
- Testing: 20%
- Description: Assigns random RGB values to the masked regions.
- Results:
Metric Value Mean Squared Error (MSE) 1139.2030 Mean Absolute Error (MAE) 7.0292 - Visualization:

- Description: Splits the masked region into 4 subregions and predicts pixel values using linear regression for each subregion.
- Training:
- Weights are initialized using Kaiming Initialization.
- Gradient descent is used to train the weights.
- Results:
Metric Value Mean Squared Error (MSE) 446.1911 Mean Absolute Error (MAE) 4.4127 - Visualization:

- Description: A CNN that predicts the mean RGB value of subregions in the masked areas.
- Training:
- Detects masked regions and splits them into (4 imes 4) subregions.
- Uses Mean Squared Error (MSE) loss for training.
- Results:
Metric Value Mean Squared Error (MSE) 336.2765 Mean Absolute Error (MAE) 3.4692 - Visualization:
- Description: A CNN with spatial and channel attention mechanisms for enhanced reconstruction.
- Components:
- Spatial Attention: Highlights important spatial regions.
- Channel Attention: Adjusts the importance of individual feature channels.
- Residual and Skip Connections: Aid in gradient flow and reuse features.
- Results:
Metric Value Mean Squared Error (MSE) 57.1340 Mean Absolute Error (MAE) 1.3617 - Visualization:
The following table summarizes the performance of all models:
| Metric | Model | Value |
|---|---|---|
| Mean Squared Error (MSE) | Baseline | 1139.2030 |
| Linear Regression | 446.1911 | |
| Neural Network | 336.2765 | |
| Attention (Best) | 57.1340 | |
| Mean Absolute Error (MAE) | Baseline | 7.0292 |
| Linear Regression | 4.4127 | |
| Neural Network | 3.4692 | |
| Attention (Best) | 1.3617 |
- Downloading: Retrieve the dataset from the Hugging Face repository.
- Dataset Structure:
dataset/ train/ image1.jpg image1_masked.jpg ... validation/ ... test/ ... - Example Dataset: Preprocessed datasets are available here.
- Format: Trained models are saved as
.pthfiles. - Example Models: Available for download here.




