Skip to content

Repository files navigation

Photo Reconstruction Project

Table of Contents

  1. Introduction
  2. Models
  3. Results Summary
  4. Dataset Preparation
  5. Saved Models

Introduction

Background and Motivation

The goal of this project is to reconstruct missing regions in images using machine learning techniques. The reconstructed images are compared with the original ground-truth images using quantitative metrics such as Mean Squared Error (MSE) and Mean Absolute Error (MAE), as well as visual examples.

Dataset and Data Preparation

  • Dataset: Curated using 10 selected classes from ImageNet: butterfly, panda, parrot, pomeranian, goldfish, elephant, monkey, Persian cat, penguin, and red panda.
  • Image Dimensions: Each image is resized to (224 imes 224) pixels.
  • Masked Regions: Blacked-out regions of random sizes ranging from (30 imes 30) to (60 imes 60) pixels simulate missing data. The positions of the masks are randomized.
  • Dataset Splits:
    • Training: 70%
    • Validation: 10%
    • Testing: 20%

Models

Baseline Model

  • Description: Assigns random RGB values to the masked regions.
  • Results:
    Metric Value
    Mean Squared Error (MSE) 1139.2030
    Mean Absolute Error (MAE) 7.0292
  • Visualization: Baseline Model Output

Linear Regression Model

  • Description: Splits the masked region into 4 subregions and predicts pixel values using linear regression for each subregion.
  • Training:
  • Results:
    Metric Value
    Mean Squared Error (MSE) 446.1911
    Mean Absolute Error (MAE) 4.4127
  • Visualization: Linear Regression Model Output

Basic Neural Network

  • Description: A CNN that predicts the mean RGB value of subregions in the masked areas.
  • Training:
    • Detects masked regions and splits them into (4 imes 4) subregions.
    • Uses Mean Squared Error (MSE) loss for training.
  • Results:
    Metric Value
    Mean Squared Error (MSE) 336.2765
    Mean Absolute Error (MAE) 3.4692
  • Visualization:
    • Loss Curve:

      NN Loss Curve

    • Reconstruction Quality: NN Reconstruction Output


Attention Model

  • Description: A CNN with spatial and channel attention mechanisms for enhanced reconstruction.
  • Components:
    • Spatial Attention: Highlights important spatial regions.
    • Channel Attention: Adjusts the importance of individual feature channels.
    • Residual and Skip Connections: Aid in gradient flow and reuse features.
  • Results:
    Metric Value
    Mean Squared Error (MSE) 57.1340
    Mean Absolute Error (MAE) 1.3617
  • Visualization:
    • Architecture: Attention Model Architecture
    • Loss Curve: Attention Model Loss Curve
    • Reconstruction Quality: Attention Model Reconstruction Output

Results Summary

The following table summarizes the performance of all models:

Metric Model Value
Mean Squared Error (MSE) Baseline 1139.2030
Linear Regression 446.1911
Neural Network 336.2765
Attention (Best) 57.1340
Mean Absolute Error (MAE) Baseline 7.0292
Linear Regression 4.4127
Neural Network 3.4692
Attention (Best) 1.3617

Dataset Preparation

  • Downloading: Retrieve the dataset from the Hugging Face repository.
  • Dataset Structure:
    dataset/
      train/
        image1.jpg
        image1_masked.jpg
        ...
      validation/
        ...
      test/
        ...
    
  • Example Dataset: Preprocessed datasets are available here.

Saved Models

  • Format: Trained models are saved as .pth files.
  • Example Models: Available for download here.

About

No description, website, or topics provided.

Resources

Stars

8 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages