The Customer Segmentation Data Warehouse (DW) project is designed to analyze customer spending behavior and segment customers into meaningful groups using data warehousing and machine learning techniques.
This project follows a complete ETL → Preprocessing → Dimensionality Reduction → Clustering → Prediction pipeline and finally allows the user to input spending amount and time (in days) to predict the customer segment.
- Python 3
- Pandas & NumPy – Data manipulation
- Scikit-learn – PCA & K-Means clustering
- Matplotlib / Seaborn – Data visualization
- Data Warehouse Concepts – ETL, analytical processing
customer_segmentation_dw/
│
├── data/ # Raw and processed datasets
├── scripts/ # Python scripts for each pipeline stage
│ ├── etl.py
│ ├── preprocessing.py
│ ├── pca.py
│ ├── clustering.py
│ └── user_segmentation.py
│
├── requirements.txt # Required Python libraries
├── README.md # Project documentation
└── .gitignore
Make sure Python 3 is installed, then run:
pip install -r requirements.txtpython3 scripts/etl.py📌 This step:
- Extracts raw customer data
- Cleans and transforms it
- Loads it into a structured analytical format
python3 scripts/preprocessing.py📌 This step:
- Handles missing values
- Applies feature scaling (standardization)
- Prepares data for PCA and clustering
python3 scripts/pca.py📌 This step:
- Reduces high-dimensional data
- Preserves maximum variance
- Improves clustering performance and visualization
python3 scripts/clustering.py📊 Output:
- A customer clustering graph will appear on the screen
python3 scripts/user_segmentation.py📌 This step:
-
Asks the user to input:
- Amount spent
- Time period (in days)
🧠 Based on the trained model, the system:
- Predicts the customer segment
- Displays the customer category on the terminal
Enter amount spent: 12000
Enter time in days: 45
Predicted Customer Segment: High-Value Customer
- Data Warehouse approach ensures structured analytical processing
- Standardization is used instead of normalization for PCA compatibility
- PCA reduces dimensionality while retaining variance
- K-Means clustering groups customers based on spending behavior
- User-driven prediction simulates real-world business decision making
- Customer behavior analysis
- Marketing strategy optimization
- Business intelligence systems
- Decision support systems
Shivam Gautam BCA – Data Warehousing & Analytics Project
This project is for academic and learning purposes only.